What happens to fine-tuned models as data and base models change?
Lugman Hussain Khan
Specialised models can make narrow AI use cases significantly more practical to deploy. Instead of relying on increasingly large general-purpose models, a smaller model can be trained around a well-defined task, evaluated against the data that matters to the application, and optimised for the accuracy, cost, and latency requirements of the product.
The development process usually follows a familiar path. Understand the use case, collect representative data, establish an evaluation set, select a suitable base model, fine-tune it, and validate the result.
It is tempting to treat training as the final step. In practice, it rarely is. Once a specialised model is deployed, two things continue to change:
- The composition of the input data changes.
- The capabilities of the underlying base models improve.
Either change can create a reason to revisit a model that was previously considered complete. Fine-tuning is therefore better thought of as part of an ongoing model lifecycle rather than a one-time optimisation exercise.

There are exceptions. If the input distribution is genuinely fixed and model accuracy has already saturated, repeated retraining may have little value. But many production systems do not operate under those conditions.
We wanted to understand what these changes look like in a concrete use case.
Testing the model lifecycle on invoice extraction
In our earlier InvoiceIQ benchmark, we showed that a fine-tuned model in the 8B parameter range could reach accuracy comparable to much larger frontier models while operating at a substantially lower inference cost.
Invoice extraction is also a useful example of why model maintenance matters. A company may onboard new suppliers. Existing suppliers may redesign their invoices. Scanned documents may be replaced by digitally generated PDFs. New geographies can introduce different conventions. Internal workflows can change which documents reach the extraction system.
At the same time, the model ecosystem continues to move forward. We ran two experiments to study these effects independently.
How much does a better base model help?
Model labs regularly release new generations with better general capabilities. We trained three generations from the same Qwen model family at a similar parameter scale using the same invoice dataset and training setup.
The base models improved substantially across generations. Accuracy increases from 38.26% to 50.20%, a gain of almost 12 percentage points.

Fine-tuned accuracy, meanwhile, stays within a relatively narrow range of 78.54% to 80.16%. The category-level results make this clearer.
| Fine-tuned model | Handwritten | Scanned | Digital |
|---|---|---|---|
| Qwen2.5-VL-7B | 61.90% | 81.78% | 95.00% |
| Qwen3-VL-8B | 63.27% | 82.59% | 99.00% |
| Qwen3.5-9B | 63.27% | 82.59% | 96.00% |
After fine-tuning, handwritten and scanned document performance is almost identical across the newer models. Digital invoices are already close to saturation. So a better base model does not automatically produce an equivalent improvement after task-specific training.
Base-model improvements can still translate to fine-tuned models, particularly when the fine-tuning dataset is relatively small or when the improvement addresses an architectural bottleneck. A new base release is worth evaluating, but its benchmark improvements should not be treated as a proxy for the gains a fine-tuned application will receive.
How data drift affects specialised models
A specialised model only reflects the environment it was trained on. In invoice extraction, this means the training data captured a particular set of vendors, page layouts, and incoming document types on a specific date. Because everyday business operations evolve, that mix never stays the same for long.
New suppliers get added. Existing suppliers change their invoice templates. The volume coming from individual suppliers changes. To simulate this evolution, we created three cumulative snapshots of the invoice dataset:
- invoices available through 2024
- invoices available through 2025
- invoices available through 2026
We trained the same Qwen3-VL-8B-Instruct model for each snapshot, using the same training configuration. The main difference between the models was the composition of the training data.
| Training data | 2024 invoices | 2025 invoices | 2026 invoices |
|---|---|---|---|
| Qwen3-VL-8B-Instruct | 46% | 51% | 55% |
| Model 2024 | 76% | 58% | 52% |
| Model 2025 | 75% | 79% | 72% |
| Model 2026 | 78% | 81% | 86% |
Model 2024 performs well on the invoices it was trained on, reaching 76% on the 2024 group. On newer invoices, its accuracy drops to 58% on 2025 invoices and 52% on 2026 invoices, which is below the base model without any fine-tuning (55%).
Adding the 2025 data closes much of that gap. On the newest invoices, accuracy improves from 52% to 72%. With the complete 2026 snapshot, it improves again to 86%. That is a 34 percentage point difference between the earliest and latest training snapshots on the newest invoice distribution.
At the same time, adding newer data does not meaningfully hurt performance on older invoices. Accuracy on the 2024 group remains within a narrow range of 75 to 78 percent across the three models.
A model does not lose capability over time. Instead, its training data covers an increasingly smaller slice of the real environment. In production, this drift creates multiple points of failure.
The model struggles with newer invoice formats and unfamiliar suppliers, while also getting stuck on familiar dates. For instance, a model trained through 2024 frequently misdates invoices from 2025 and 2026, writing 2024 or earlier simply because it leans toward the years it saw during training.
Managing catastrophic forgetting
Once new data becomes available, a natural approach is to continue training the existing model only on those new examples. Our experiment shows why this needs some care.
We took Model 2024 and continued training it on 2025 documents only. Then we compared it with Model 2025, which trained on the same 2025 documents plus everything before them.

The newer-only model retained its ability to process the data it had just seen, but lost a large part of its performance on earlier distributions. This is a useful example of how narrow retraining can introduce catastrophic forgetting.
The solution is fairly simple. When updating a specialised model, the training mixture should include both new examples and representative samples from previous distributions. For a small dataset, retraining from the base model using the updated cumulative dataset can also be the simpler option.
The takeaway
Fine-tuning is not a one-time step that ends when a model reaches a target accuracy. A specialised model is tied to both the capabilities of its base model and the distribution of data it was trained on.
Our experiments show that newer base models are worth evaluating, but improvements at the base level do not always translate proportionally after fine-tuning. Changes in the production data can have a much more direct effect. As new inputs, workflows, user behaviours, and system dependencies emerge, a model trained on an older snapshot can gradually become less representative of the environment it is expected to operate in.
This makes maintaining the dataset and evaluation set just as important as maintaining the model itself. The goal is to keep evaluating whether the model still represents the task it is expected to solve, and to update it when the data or the underlying model capabilities create a meaningful reason to do so.