Blog
engineering

What happens to fine-tuned models as data and base models change?

L

Lugman Hussain Khan

Specialised models can make narrow AI use cases significantly more practical to deploy. Instead of relying on increasingly large general-purpose models, a smaller model can be trained around a well-defined task, evaluated against the data that matters to the application, and optimised for the accuracy, cost, and latency requirements of the product.

The development process usually follows a familiar path. Understand the use case, collect representative data, establish an evaluation set, select a suitable base model, fine-tune it, and validate the result.

It is tempting to treat training as the final step. In practice, it rarely is. Once a specialised model is deployed, two things continue to change:

  1. The composition of the input data changes.
  2. The capabilities of the underlying base models improve.

Either change can create a reason to revisit a model that was previously considered complete. Fine-tuning is therefore better thought of as part of an ongoing model lifecycle rather than a one-time optimisation exercise.

Fine-tuned model lifecycle loop

There are exceptions. If the input distribution is genuinely fixed and model accuracy has already saturated, repeated retraining may have little value. But many production systems do not operate under those conditions.

We wanted to understand what these changes look like in a concrete use case.

Testing the model lifecycle on invoice extraction

In our earlier InvoiceIQ benchmark, we showed that a fine-tuned model in the 8B parameter range could reach accuracy comparable to much larger frontier models while operating at a substantially lower inference cost.

Invoice extraction is also a useful example of why model maintenance matters. A company may onboard new suppliers. Existing suppliers may redesign their invoices. Scanned documents may be replaced by digitally generated PDFs. New geographies can introduce different conventions. Internal workflows can change which documents reach the extraction system.

At the same time, the model ecosystem continues to move forward. We ran two experiments to study these effects independently.

How much does a better base model help?

Model labs regularly release new generations with better general capabilities. We trained three generations from the same Qwen model family at a similar parameter scale using the same invoice dataset and training setup.

The base models improved substantially across generations. Accuracy increases from 38.26% to 50.20%, a gain of almost 12 percentage points.

Base vs fine-tuned accuracy by model generation

Fine-tuned accuracy, meanwhile, stays within a relatively narrow range of 78.54% to 80.16%. The category-level results make this clearer.

Fine-tuned modelHandwrittenScannedDigital
Qwen2.5-VL-7B61.90%81.78%95.00%
Qwen3-VL-8B63.27%82.59%99.00%
Qwen3.5-9B63.27%82.59%96.00%

After fine-tuning, handwritten and scanned document performance is almost identical across the newer models. Digital invoices are already close to saturation. So a better base model does not automatically produce an equivalent improvement after task-specific training.

Base-model improvements can still translate to fine-tuned models, particularly when the fine-tuning dataset is relatively small or when the improvement addresses an architectural bottleneck. A new base release is worth evaluating, but its benchmark improvements should not be treated as a proxy for the gains a fine-tuned application will receive.

How data drift affects specialised models

A specialised model only reflects the environment it was trained on. In invoice extraction, this means the training data captured a particular set of vendors, page layouts, and incoming document types on a specific date. Because everyday business operations evolve, that mix never stays the same for long.

New suppliers get added. Existing suppliers change their invoice templates. The volume coming from individual suppliers changes. To simulate this evolution, we created three cumulative snapshots of the invoice dataset:

We trained the same Qwen3-VL-8B-Instruct model for each snapshot, using the same training configuration. The main difference between the models was the composition of the training data.

Data drift: accuracy by training snapshot

Training data2024 invoices2025 invoices2026 invoices
Qwen3-VL-8B-Instruct46%51%55%
Model 202476%58%52%
Model 202575%79%72%
Model 202678%81%86%

Model 2024 performs well on the invoices it was trained on, reaching 76% on the 2024 group. On newer invoices, its accuracy drops to 58% on 2025 invoices and 52% on 2026 invoices, which is below the base model without any fine-tuning (55%).

Adding the 2025 data closes much of that gap. On the newest invoices, accuracy improves from 52% to 72%. With the complete 2026 snapshot, it improves again to 86%. That is a 34 percentage point difference between the earliest and latest training snapshots on the newest invoice distribution.

At the same time, adding newer data does not meaningfully hurt performance on older invoices. Accuracy on the 2024 group remains within a narrow range of 75 to 78 percent across the three models.

A model does not lose capability over time. Instead, its training data covers an increasingly smaller slice of the real environment. In production, this drift creates multiple points of failure.

The model struggles with newer invoice formats and unfamiliar suppliers, while also getting stuck on familiar dates. For instance, a model trained through 2024 frequently misdates invoices from 2025 and 2026, writing 2024 or earlier simply because it leans toward the years it saw during training.

Managing catastrophic forgetting

Once new data becomes available, a natural approach is to continue training the existing model only on those new examples. Our experiment shows why this needs some care.

We took Model 2024 and continued training it on 2025 documents only. Then we compared it with Model 2025, which trained on the same 2025 documents plus everything before them.

Catastrophic forgetting: cumulative vs new-only training

The newer-only model retained its ability to process the data it had just seen, but lost a large part of its performance on earlier distributions. This is a useful example of how narrow retraining can introduce catastrophic forgetting.

The solution is fairly simple. When updating a specialised model, the training mixture should include both new examples and representative samples from previous distributions. For a small dataset, retraining from the base model using the updated cumulative dataset can also be the simpler option.

The takeaway

Fine-tuning is not a one-time step that ends when a model reaches a target accuracy. A specialised model is tied to both the capabilities of its base model and the distribution of data it was trained on.

Our experiments show that newer base models are worth evaluating, but improvements at the base level do not always translate proportionally after fine-tuning. Changes in the production data can have a much more direct effect. As new inputs, workflows, user behaviours, and system dependencies emerge, a model trained on an older snapshot can gradually become less representative of the environment it is expected to operate in.

This makes maintaining the dataset and evaluation set just as important as maintaining the model itself. The goal is to keep evaluating whether the model still represents the task it is expected to solve, and to update it when the data or the underlying model capabilities create a meaningful reason to do so.

Ready to make AI part of how your business operates?

Let's identify the workflows where AI can create the greatest value and determine the right way to build them.