Skip to main content

Command Palette

Search for a command to run...

What Does Fine-Tuning Actually Learn? Testing QLoRA with Deliberately Controlled Data

Updated
•7 min read•View as Markdown
What Does Fine-Tuning Actually Learn? Testing QLoRA with Deliberately Controlled Data

Fine-tuning is often treated as a way to "embed knowledge" into a model — but what actually happens under the hood can be far more layered than that. This post builds a small, controlled experiment — two fictional customer-service businesses, a training set of known size, and three different ways of testing the answers — to see directly what actually gets locked in, and what only looks like it did.

💡 How to read this post: there are a few "Guess First" boxes scattered through the text. Take a guess before checking the answer.


Table of Contents

  1. Why a Controlled Experiment?

  2. Three Ways to Test One Answer

  3. From Raw Text to Chat Model: Where Fine-Tuning Fits

  4. Plot Twist: Format Locks In First, Facts Come Later

  5. Brand Drift: Real, But Not About "Can't"

  6. Cheat Sheet

  7. Test Your Understanding


Why a Controlled Experiment?

Fine-tune a customer-service chatbot, see decent-looking answers come out, and it's tempting to conclude "the model learned it." But learned what, exactly? The greeting format? The brand name? The actual facts about its policies? All three can look equally "correct" on the surface while being learned at very different speeds and in very different ways.

This experiment is built specifically to tell them apart: two fictional domains — GadaiKita (a pawnshop) and TeknoMart (an electronics retailer) — each with 10 facts × 6 phrasings (60 training examples), trained at three separate checkpoints: 0 epochs (untrained baseline), 3 epochs, and 15 epochs.

Three Ways to Test One Answer

Each domain gets tested with three sets of questions, each measuring something different:

Set Contents What it tests
train_exact 10 questions identical to training data Memorization
paraphrase 10 questions, same meaning, new wording Generalization to new phrasing
heldout 4 questions about facts never in training Behavior when a fact is unknown

Every answer is checked automatically for four things: format compliance (correct greeting and closing), correct brand name, presence of key fact keywords, and repetition. Nothing gets scored by manual reading — it's all in code.

🎯 Guess First: At epoch 3, what percentage of train_exact answers — questions identical to the training data — do you think already have the correct format (exact greeting + closing)?

See the answer

70%, in both domains. But that's only half the story — skip to the Plot Twist section below to see how many of the facts were right at that same point.

From Raw Text to Chat Model: Where Fine-Tuning Fits

Building a chat model usually goes through three stages: pretraining on raw, unlabeled text (producing a base model with basic capabilities like text completion), supervised fine-tuning (SFT) on instruction-answer pairs (producing a model that follows instructions), and preference alignment on human/AI preference data (producing the final, more aligned chat model).

The experiment in this post lives entirely in the SFT stage — training a base model directly on customer-service question-answer pairs, with no alignment step afterward. That matters for setting expectations: SFT is good at teaching the shape of a response, but as the next section shows, that doesn't automatically mean the content gets learned at the same pace.

Plot Twist: Format Locks In First, Facts Come Later

At epoch 3, both domains already answer with the correct greeting and closing 70% of the time — even on questions identical to the training data. But the facts in those same answers?

0%.

Format learned faster than facts

Facts only catch up at epoch 15 — 100% on train_exact. But the moment the question gets paraphrased (same meaning, different wording), that number drops back down to 40% (GadaiKita) and 50% (TeknoMart). The model can greet perfectly and repeat a fact it has literally memorized, but generalizing that fact to a question it hasn't seen before is a much harder ask than generalizing the format.

This lines up with how SFT works: the model picks up response patterns (sentence structure, greeting, closing) much faster than fact associations, which need more exposure and variation before they generalize.

Brand Drift: Real, But Not About "Can't"

One more finding: at epoch 3, GadaiKita already gets its own name right 100% of the time. TeknoMart doesn't — some answers instead say "TeknoSupport" or "TeknoMart Support," names that are close but wrong.

Brand drift by epoch

The tempting conclusion here is "fine-tuning on small data can't lock in literal entities." But once training continues to epoch 15, the drift disappears entirely — TeknoMart also reaches 100%. So it's not about "can't," it's about not enough exposure yet. TeknoMart needed more steps to lock in its own name than GadaiKita did — likely because "TeknoMart" has more similar-sounding variants that could confuse the model early on (TeknoExpress, TeknoBank, and so on) compared to the more distinctive "GadaiKita."

There's one more wrinkle that shows up right at epoch 15: TeknoMart's format_pct on heldout questions actually drops from 75% to 50%, even as brand_pct on the same set climbs to 100%. Training loss in both domains is very low at this point (0.33–0.34) — a sign of overfitting on just 60 examples per domain. The symptom doesn't always show up in the same metric: one GadaiKita answer at epoch 15 contains a nonsensical phrase ("...only applies during word-time..." — gibberish in the original Indonesian), even though GadaiKita's own format and brand metrics at that point look perfect (100%/100%). Heavier training locks the brand name in perfectly, but breaks sentence-level coherence on genuinely new prompts — sometimes visible as a metric dropping (TeknoMart), sometimes only visible if you actually read the answer (GadaiKita).

Cheat Sheet

  • train_exact vs paraphrase vs heldout — three angles that measure different things: memorization, generalization to new phrasing, and behavior on unknown facts.

  • SFT learns format (style, response structure) much faster than facts — don't mistake "the response looks polished" for "the content is correct."

  • Entity drift (getting a name wrong) in short training runs isn't necessarily a permanent failure — it might just need more steps.

  • Training too long on small data can lock in one thing (a brand name) while breaking another (sentence coherence) — overfitting doesn't fix everything at once.

  • Always add an automated check for data leakage between training and evaluation sets — a biased baseline (an untrained model that somehow looks like it "knows something") is usually a sign the evaluation questions are flawed, not that the model is smart.

Test Your Understanding

1. If a customer-service model greets perfectly and its answer format looks clean, does that prove the content of the answer is also correct?

No. This experiment shows format can reach 70–100% well before facts catch up — even on questions identical to training data, facts can sit at 0% at the same training point where format is nearly perfect.

2. Why did TeknoMart need more epochs than GadaiKita to stop misnaming itself?

Not because SFT is fundamentally incapable of locking in entities — once training continued to epoch 15, the drift disappeared completely in both domains. The likely reason is that "TeknoMart" has more similar-sounding variants that could confuse the model during short training, compared to the more distinctive "GadaiKita."

3. If longer training fixed the brand drift, why doesn't this post conclude "more training is always better"?

Because at that same training point (epoch 15), a different side effect showed up: format_pct on heldout questions actually dropped, and one answer became linguistically incoherent. The very low training loss at this point points to overfitting — fixing one metric can break another, so "longer" isn't an automatic win.


The full notebook (with an automated data-leakage guard and real run output intact) is at the GitHub repo. The pretraining-SFT-alignment pipeline mentioned here is covered in more depth in the post on fine-tuning Llama 3.1 8B.

AI Engineering Study Notes

Part 1 of 50

A personal collection of AI engineering study notes — covering computer vision, deep learning, and model deployment — built from AI Super Class coursework and independent exploration.

More from this blog

S

Shaka's AI Journal

74 posts

A personal AI engineering journal — documenting hands-on learning in computer vision, deep learning, data pipelines, and model deployment. Study notes, working code, and honest write-ups from coursework and independent projects, published in Indonesian and English.