A Bigger Model Didn't Fix My Kaggle Project
A beginner-friendly knee-MRI experiment journal: the problem, model choices, AUC, failed hypotheses, and US$75 plus C$155 spent learning what actually helps.
I thought improving an image model would be a fairly direct climb: start small, use a bigger network, improve its training answers, then combine the best runs.
Instead, my “better” model won twice and failed badly on a repeat. Newer training labels made one comparison worse. An ensemble's promising gain shrank on a fresh test.
This is what happened while learning through a knee-MRI competition—and the ML ideas that make those failures useful. You don't need to have read the earlier posts.
What are we trying to build?
Kaggle runs competitions where participants train on supplied data, then predict answers for unseen test examples. In this RSNA competition, the input is a knee MRI examination; the output is a probability for each of 12 findings, including ligament injuries, meniscus damage, arthritis and fluid-related findings.
This is multi-label classification: several findings can coexist. We aren't choosing one diagnosis from a menu. We produce 12 separate scores between 0 and 1.
An MRI examination isn't one photograph. It contains stacks of slices, often viewed from the side, front and across the knee—sagittal, coronal and axial planes. Different scan sequences also emphasize different tissues. Sampling the wrong slices can hide useful evidence before the network ever sees it.
The second challenge is the answer sheet. Thousands of studies had written radiology reports, but our expert-labelled image-reference set was small. We converted reports into weak labels: imperfect training hints. “Not mentioned” isn't necessarily “absent,” so uncertain or unsupported answers could be downweighted or masked, meaning ignored by the training loss rather than treated as healthy.
Language models helped read the reports. They were not our MRI prediction models. At test time, our image network had to predict from scans, without a report supplying the answer.
Why these models—not just the newest one?
We started with transfer learning: reuse an image network already trained elsewhere, then adapt it to our task. The backbone, or encoder, turns images into visual features; a smaller prediction head converts those features into our 12 scores. Fine-tuning updates some or all of those learned weights.
That gives us a useful starting point, not medical expertise. Recognizing patterns in ordinary photographs doesn't guarantee recognition of a small ligament injury.
| Model | What it brings | Why we tested it |
|---|---|---|
| ConvNeXt-Tiny | An efficient convolutional image network: learns visual patterns through local image operations. | An affordable baseline to make the pipeline work and give every challenger a reference. |
| CoAtNet | Combines convolution with attention, which can connect information across image regions. | A plausible richer image reader, backed by a promising public recipe for this competition. |
| DINOv2 | Reusable visual features learned through self-supervised pretraining, without manual class labels. | Could a small head use those existing features without expensive full-network adaptation? |
| OrthoFoundation | A model pretrained on knee imaging, including MRI slices and X-rays. | Would domain-specific pretraining transfer better than general-image pretraining? |
For the last two, our initial tests froze the encoder and trained a head. That is cheaper and isolates whether existing features are useful. It is not equivalent to full fine-tuning; a failed frozen-feature test doesn't prove the whole model family is unsuitable.
We also tested how to combine slices. Mean pooling averages their features; attention pooling learns which slices deserve more influence. More useful image coverage and a bigger network are separate hypotheses—not interchangeable upgrades.
How do we know a change really helped?
A baseline is the reference recipe. A matched comparison changes one main ingredient while keeping the data split, inputs and training setup aligned. Otherwise, we might credit a model for an easier test set.
We split eligible studies into five folds. For a run, one fold stays out of training: a practice exam the model cannot study directly. A different fold tests a different held-out group. A seed controls random initialization and sampling; repeating the experiment with a different seed tests sensitivity to those random choices.
All slices from an examination stay together; otherwise, neighboring pictures of the same knee could appear in training and validation. We also group duplicate reports. These are guards against data leakage: letting information from the test influence learning and make the score deceptively good.
Before a long run, a smoke test checks a tiny workload: do images load, gradients stay finite, weights update, and saved predictions replay? Passing it means the machinery works—not that the model is accurate.
During training, loss is the mistake penalty used to update weights through backpropagation. Our binary cross-entropy loss penalizes confident wrong answers. Validation loss applies that penalty to held-out examples, without updating weights from them; we use it to select a saved checkpoint.
An epoch is a pass through the training examples, not a turn devoted to one condition. All 12 outputs are trained across epochs. More passes can improve learning, but can also encourage overfitting: learning peculiarities of the training data instead of transferable evidence.
AUC, explained with four comparisons
Accuracy asks how many yes/no decisions were correct at a chosen cutoff. That can mislead: if 99 of 100 cases are negative, “always say no” gets 99% accuracy while finding no positives.
ROC-AUC asks a different question: does a positive example receive a higher score than a negative one? Its name means “area under the receiver operating characteristic curve,” but you can understand it through ranking. Google's ML course explains this interpretation.
Imagine two positive cases scored 0.8 and 0.4, and two negatives scored 0.6 and 0.2. There are four positive–negative pairs. The 0.8 positive outranks both negatives; the 0.4 positive outranks only the 0.2 negative. Three correct rankings out of four gives AUC 0.75. Ties receive half credit; chance-level ranking is 0.5, perfect ranking is 1.
This competition averages the AUCs of the 12 conditions: macro AUC. Each condition contributes equally, even if one is much rarer. AUC doesn't establish calibration—whether cases scored 0.8 are actually positive about 80% of the time—or whether a clinical decision is safe.
An AUC of 0.834 is not “83.4% of diagnoses correct.” And training loss is not AUC: a lower mistake penalty need not improve the ordering we ultimately measure.
There is another trap: different answer sheets. Our local report-label tests, small expert-labelled development set and public leaderboard measure different cohorts and labels. The public leaderboard is also only a provisional view of the hidden test evaluation. Scores across these settings are not a single progress curve.
Our earlier ConvNeXt submission scored 0.811 public AUC; a saved CoAtNet scored 0.834. That CoAtNet's local report-label score was 0.905, not a promise of 0.905 on the leaderboard. An external public CoAtNet recipe reported 0.924—its author's result, not ours.
Three assumptions the experiments challenged
1. “A stronger backbone will reliably win.” CoAtNet beat matched ConvNeXt controls on two report-label comparisons: 0.905 vs 0.873, then 0.886 vs 0.857. Repeating the second comparison with a different seed reversed the story: 0.733 vs 0.866. We kept the useful checkpoint, but rejected the claim of a robustly better recipe.
2. “A stronger language model means better training labels.” We revised report extraction with GPT-6 Sol after an earlier small-model/escalation pipeline. All 4,349 revised reports passed structural checks after one repair. Yet a matched image-model comparison on 31 previously consulted expert-labelled development cases fell from 0.789 with older labels to 0.730 with revised labels. This reused, small set is diagnostic—not an independent final test. A separate 27-case audit remained unopened. Correctly formatted answers and answers aligned with the competition are different achievements.
3. “Combining models will cancel out their errors.” An ensemble averages predictions; it helps when useful strengths complement one another, not just because there are more models. A fixed 50/50 blend improved two discovery runs by 0.0099 and 0.0073. On a fresh fold, its mean gain shrank to 0.0026, below our preset 0.005 requirement. Not zero benefit, but insufficient evidence to promote it. We didn't retune weights on that fresh fold to rescue the story.
Keeping the best run is useful engineering. Reporting only the best run is bad science.
The money—and the accounting gaps
Here is what I reported spending through October 2, 2026. These are project expenses, not current provider price quotes or precise per-experiment invoices.
| Service | Reported expense | What it bought |
|---|---|---|
| OpenAI API for labelling | US$50 | Report extraction, review and revised-label experiments. The revised-label portion is estimated at US$14.30 from token usage; included, not extra. |
| Runpod | US$25 | Paid GPU training/infrastructure, including the overnight monitoring failure. Exact per-fit and idle-time allocation is unknown. |
| Kaggle | US$0 paid cash | Included GPU/TPU quota plus CPU notebooks. Quota is limited; total historical accelerator hours were not reconciled. |
| API + Runpod subtotal | US$75 | Reported spending, not a precise per-experiment usage invoice. |
| ChatGPT subscription | C$155 | Coding/research assistance, separate from the labelling API bill. Shared tooling rather than a per-training-run charge. |
Total reported outlay: US$75 + C$155. Currencies stay separate. Human time, local electricity and an invented dollar value for Kaggle quota aren't included.
Two measured Kaggle feature-extraction stages took 2h 44m 35s and 2h 38m 51s. The latest first-seed training runs took roughly 1h 52m and 2h 6m, excluding preparation and queue time. These are individual stage/fit runtimes, not a reconciled total of accelerator-quota hours.
The avoidable mistake: a rented GPU kept running while my sleeping laptop stopped coordinating follow-up work. Credits ran out. We recovered useful artifacts, but cannot assign an exact idle-time cost. Cleanup must run remotely; checkpoints must survive the machine.
Final experiment ledger: hypothesis, result, cost, lesson
The remaining tests ask whether more detail, more useful views or more supervision will help. An auxiliary task adds a related question during training—for example, a broader finding alongside a strict injury definition. A local crop zooms into anatomy rather than shrinking the entire image.
This is the compact audit trail, not recommended settings. K = included Kaggle quota/local compute, US$0 additional cloud cash. R = part of the shared US$25 Runpod spending, not another US$25 per row. The subscription supports work across rows. Unqualified AUCs are report-label validation, not leaderboard scores; two numbers describe repeated runs unless otherwise stated.
| Experiment | Hypothesis | Observed result | Cost / compute | What we learned |
|---|---|---|---|---|
| Three MRI planes vs one | More viewpoints expose useful anatomy. | 0.735 → 0.789, consulted 31-case image-development set. | K | Useful exploratory signal. |
| Condition-specific plane weights | Each condition needs a different view. | 0.789 → 0.786, same development set. | K | Static weights didn't help; learned specialists remain a different hypothesis. |
| Fine-tuned CoAtNet | Richer backbone beats ConvNeXt. | Two wins, then 0.733 vs 0.866; saved model 0.834 public. | R + K | Checkpoint useful; recipe unreliable. |
| Mean pooling | Simpler aggregation repairs instability. | 0.542 vs 0.733 with attention, same seed. | R | Not a repair. |
| Frozen CoAtNet / head warm-start | Preserve features; learn the head first. | Frozen 0.793/0.798; warm-start 0.808/0.810; controls 0.857/0.866. | K | Consistency isn't enough. |
| Frozen DINOv2 / OrthoFoundation | Reuse stronger or knee-specific features. | DINOv2 mean 0.816, attention 0.788; OrthoFoundation 0.752/0.763. Neither promoted. | K | Full fine-tuning remains untested by these screens. |
| Revised report labels | Stronger text reader improves supervision. | 0.789 → 0.730, consulted 31-case image-development comparison. | US$14.30 estimate, within US$50 API | Valid formatting doesn't establish alignment. |
| Auxiliary task | Related broad findings help strict targets. | 0.856/0.865 vs controls 0.857/0.866. | K | Extra supervision didn't help here. |
| Fixed ensemble | Complementary errors cancel out. | Fresh-fold mean gain +0.0026, below preset +0.005. | K | Discovery gain wasn't strong enough on confirmation. |
| Higher resolution | More pixels reveal small abnormalities. | 384 pixels 0.852 vs 224 pixels 0.854, first test. | K | Bigger inputs alone didn't help. |
| More scan types | Broader coverage helps at fixed image budget. | Seed deltas +0.0035/−0.0071. | K | Broader coverage can sacrifice useful detail. |
| Local-detail crops | Zooming preserves small structures. | Pixel checks passed; cache/time limits blocked training. | K + CPU | Feasibility failure, not evidence against accuracy. |
| TPU / recovery checks | Better execution reduces wasted work. | Synthetic TPU and GPU recovery checks passed; no TPU MRI accuracy result. | K | Reliability isn't model quality. |
| Balanced loss | Give rare conditions/classes more training influence. | First seed: control 0.9031, balanced 0.8941; repeat pending. | K | Preliminary decline; not a stronger candidate. |
Why training isn't a straight path
Think of preparing for an exam. A larger model is a more capable student. More epochs are more passes through the notes. But neither fixes misleading notes, missing diagrams, or practice questions that differ from the actual exam. Studying longer can even teach the wrong shortcuts more thoroughly.
The practical recipe for a student is hypothesis → small systems test → matched comparison → repeat seed → fresh-fold confirmation → decision. Write the success threshold before seeing the result. Record failures and compute, not just the best score. Leave genuinely untouched data for a final check.
Our latest hypothesis rebalances the loss because equal weighting in the competition metric doesn't imply equal influence during training. The first result went backwards; the registered repeat will help characterize that result, not turn it into a victory by selective reporting. We haven't solved the competition.
That's why training is experimental science, not a shopping list of stronger models. A failed experiment is useful when it changes what you do next—and when it rules out only what it actually tested.
Part 3 of Learning Model Training. Part 2 explains the MRI, label and validation foundations. Results are a snapshot of October 2, 2026; this post stands alone. This is competition research, not a clinical diagnostic system. No patient images, reports or case identifiers are published here.
Before Training a Model, I Had to Build a Data Factory
The data, labels and validation decisions behind these experiments.