Findings-aware retrieval and reinforcement-tuned generation for free-text reporting of non-contrast head CT.
Team IUCompPath - Sanket Kachole, Spyridon Bakas (Division of Computational Pathology, Indiana University School of Medicine). Contact: Spyridon Bakas (spbakas@iu.edu).
Challenge: https://headline26.grand-challenge.org/ · Data: https://headline.ccibonn.ai
A 3D CNN is trained on weak labels mined from the reports (haemorrhage, infarct, midline shift) and then frozen. Its 256-d bottleneck is a finding-aware embedding used to retrieve radiologist-authored reports; its spatial feature map conditions a LoRA-adapted Qwen2.5-1.5B, further tuned with GRPO against the challenge metric. The generated report is added to the retrieved pool and the most central candidate is selected by BERTScore consensus. Full description in the paper.
Official validation: a composite score of 0.2400 (retrieval-only, 500 studies), comprising a BERTScore of 0.2634, ROUGE-L of 0.2800, METEOR of 0.1697, and BLEU-4 of 0.1064.
conda create -n headline python=3.10 && conda activate headline
pip install -r requirements.txtTraining used a single NVIDIA RTX 6000 Ada (48 GB); preprocessing ran on
Big Red 200. The inference container pins torch==2.4.0+cu121 for the Grand
Challenge runtime (driver 535 / CUDA 12.0, T4 or A10G, 20 min per batch, no
network access).
Download the weight bundle from Zenodo (Section 4), extract it into ./model,
then build and run the container. This reproduces the submitted system.
./do_build.sh # build the image
./do_test_run.sh # forward pass on the sample batch in test/
./do_save.sh # package as .tar.gz for Grand Challenge uploadinference.py reads the stacked 4D volumes from /input, runs
CNN → retrieve top-80 → add one generated report → BERTScore consensus, and
writes /output/radiology-text-reports.json.
Run in order. Paths are argparse defaults and should be overridden per environment.
# index the corpus and make the patient-level split
python scripts/01_inventory_training_data.py
python scripts/02_build_master_index_and_split.py
# preprocess volumes + mine findings labels (see note - HPC scripts)
# -> headline_preprocessed/volumes/*.npz
# -> findings label columns in outputs/reports_mining/master_index_clean.csv
# findings-aware encoder
python scripts/22_train_findings_classifier.py # 3D CNN, 25 epochs
python scripts/23_predict_findings_all.py # 256-d embedding per study
# retrieval + consensus -> Table 2
python scripts/24_findings_retrieval.py
python scripts/25_findings_tuning.py
# generation
python scripts/50_train_vlm.py # projector + LoRA on Qwen
python scripts/50b_train_vlm_resume.py # continue to 5 epochs total
python scripts/60_grpo.py # GRPO on the challenge metric
# ensemble + package -> Table 3
python scripts/54_ensemble_gen_retrieval.py
python scripts/56_ensemble_pool_sweep.py # picks TOP_K = 80
python scripts/55_build_vlm_bundle.py --epoch 100 --out model
# build the validation-phase submission CSV
python scripts/28_build_official_submission.pyScores are computed with scripts/headline_metric.py, a local copy of the
official Stage-1 metric (0.70 × BERTScore-F1 + 0.30 × mean(METEOR, ROUGE-L,
BLEU-4)).
Note - two preprocessing scripts are not included. Volume windowing/ resampling and findings-label mining were run on HPC; they produce
headline_preprocessed/volumes/*.npzand the label columns inmaster_index_clean.csv. Path A does not need them (the released bundle already contains the trained CNN and embeddings). For Path B, contact the authors for these two scripts or regenerate the inputs from the method in the paper.
Root:
| File | Purpose |
|---|---|
inference.py |
Container entry point (retrieve + generate) |
do_build.sh / do_save.sh / do_test_run.sh |
Build, package, test the container |
scripts/ - the reproduction pipeline:
| File | Purpose |
|---|---|
01_inventory_training_data.py |
Survey the training corpus |
02_build_master_index_and_split.py |
Master index + patient-level split |
22_train_findings_classifier.py |
Train the 3D findings CNN |
23_predict_findings_all.py |
Embed all studies with the CNN |
24_findings_retrieval.py |
Findings-aware cosine retrieval |
25_findings_tuning.py |
Tune retrieval (K, centering) |
50_train_vlm.py |
Train projector + LoRA on Qwen |
50b_train_vlm_resume.py |
Resume VLM training to 5 epochs |
60_grpo.py |
GRPO against the challenge metric |
54_ensemble_gen_retrieval.py |
Add generation to pool, consensus |
56_ensemble_pool_sweep.py |
Choose ensemble TOP_K |
55_build_vlm_bundle.py |
Package the weight bundle |
28_build_official_submission.py |
Build the submission CSV |
headline_metric.py |
Local copy of the official metric |
The remaining scripts in scripts/ are the ablation baselines and recorded
negative results behind Table 2 and the paper's discussion (numbered by
development stage). They are kept for transparency and are not needed to
reproduce the main result.
Bundle (~4.6 GB): findings_cnn.pt, projector.pt, vlm/ (merged fp16 Qwen +
LoRA), roberta-large/ (offline BERTScore).
The retrieval pool of report text (
train_reports.json, 16,468 reports) is not distributed - redistributing it would breach the data use agreement. It must be rebuilt from a separately obtained, licensed copy of the training data withscripts/30_build_model_bundle.py.
- No bone window is used, so fractures are unlikely to be detected.
- Optimisation targets text overlap, not clinical correctness; the system is not suitable for unsupervised clinical use.
- Selection is the bottleneck: 0.1935 realised against a 0.2676 oracle.
Kachole, S., Bakas, S. Findings-Aware Retrieval and Reinforcement-Tuned Generation for Free-Text Head CT Reporting. MICCAI 2026 HEADLINE Challenge.
Container scaffolding is derived from the organisers' starter kit, https://github.com/CCI-Bonn/headline26-starter-kit.