Official code for Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers
- [2026-10-10] π₯ We release the training and evaluation code of Trident, including configs, training scripts, and the Origin / Mix evaluation pipeline. Data and model checkpoints are coming soon.
- News
- Overview
- Requirements / Environment Setup
- Data Preparation
- Experiments
- Repository Structure
- Acknowledgements
- License
- Contact
- Citation
Dense retrievers perform well on text-only and image-only corpora, but real-world corpora mix text-only, image-only, and fused text-image documents. We find that retrieval performance is highly sensitive to this modality composition:
- Chaos in the Text. As image documents are progressively replaced by their text counterparts, NDCG@10 drops sharply and only partially recovers when the corpus becomes text-dominant. Irrelevant text hurts far more than an equal number of irrelevant images: irrelevant text documents are ranked above relevant images.
- Modality preference. Retrievers systematically assign higher similarity scores to text representations than to semantically equivalent image representations.
To mitigate this bias, we propose Trident:
- Modality-balanced data construction. Every document is represented by three co-equal positive views: a Text view (a detailed description), an Image view (the original image), and a Fused view (a short summary + the image).
- Multi-Positive View InfoNCE. All positive views and all in-batch negatives share a single softmax denominator, and the loss averages the log-probabilities of the positives:
Trident improves mixed-modality retrieval for both CLIP-based (jina-clip-v2) and VLM-based (Qwen3-VL-2B) retrievers, flattens the V-shaped performance surface, reduces sensitivity to text distractors, and also improves average single-modality retrieval.
The code is tested with Python 3.12, PyTorch 2.10, and transformers 4.57.1 (Qwen3-VL requires >= 4.57, and the jina-clip-v2 remote code is not yet compatible with transformers 5.x). All experiments in the paper were run on 8 Γ NVIDIA A100-80GB GPUs.
git clone https://github.com/OpenBMB/Trident.git
cd Trident
conda create -n trident python=3.12 -y
conda activate trident
pip install -r requirements.txtCheck that the environment is complete:
python -c "import sys; sys.path.insert(0, 'src'); from mmemb.utils.env import describe_env; print(describe_env())"Baseline-specific dependencies (only needed to reproduce the baseline rows):
| Backend | Extra setup |
|---|---|
unime_phi35v, visrag_ret |
pip install modelscope. VisRAG-Ret's remote code requires an older transformers (transformers==4.40.2); use a separate environment for it. |
qwen3vl_official |
Place the official qwen3_vl_embedding.py (class Qwen3VLEmbedder) from the Qwen3-VL-Embedding release in src/, or pass --qwen3vl_official_script <path>. |
Download the training data and the evaluation benchmarks:
coming soonThe Hugging Face dataset repository will be announced here when the data is released.
Expected layout:
data/
βββ Train/train.jsonl # tri-modal training data
βββ images/ # training images (paths in train.jsonl are relative to it)
βββ ChartQA/ DocVQA/ InfoVQA/ SlideVQA/ ViDoSeek/ Dude/ # Table 1
βββ Google_WIT/ MSCOCO/ VisualNews/ OVEN/ # Table 4
β βββ queries.jsonl
β βββ origin_corpus.jsonl # the dataset's original single-modality corpus -> "Origin"
β βββ mix_corpus.jsonl # mixed-modality version of the corpus -> "Mix"
β βββ qrels.jsonl
Training data. We build training instances from OpenDocVQA (41k queries) following Section 4.1: Qwen3-VL-235B-A22B-Instruct generates a detailed description and a short summary for every image document. Each line contains one query and the three positive views of its document:
{"query": [{"text": "How many stores did Saint Laurent operate in Western Europe in 2020?"}],
"positive": [{"text": "<detailed description>", "id": "d1_text"},
{"image": "chartqa/d1.png", "id": "d1_image"},
{"text": "<short summary>", "image": "chartqa/d1.png", "id": "d1_fused"}],
"doc_id": "d1", "query_id": "q1", "task": "default"}The view order (text, image, fused) is the same in every row. Optional identity fields (doc_id, query_id, positive_ids, group_id) are used to mask false negatives in the batch; missing ids fall back to content hashes. See src/mmemb/data/schema.py for all supported fields.
Evaluation data. queries.jsonl, origin_corpus.jsonl, and mix_corpus.jsonl contain one record per line, {"id": ..., "text": ..., "image": ...} (text and/or image; image paths are relative to data/<DATASET>/). qrels.jsonl contains {"query_id": ..., "doc_id": ..., "score": 1}. In the Mix setting, each corpus is converted into a mixed-modality corpus following Section 2.1 and the corpora of all datasets of a benchmark are merged into a single pool.
# Trident-Qwen3VL-2B (Qwen3-VL-2B-Instruct + LoRA)
bash scripts/train.sh configs/trident_qwen3vl.yaml
# Trident-JinaCLIP (jina-clip-v2, full-parameter)
bash scripts/train.sh configs/trident_jinaclip.yaml
# override any field from the command line
bash scripts/train.sh configs/trident_qwen3vl.yaml --set train.output_dir=outputs/my_run data.train_path=data/train/train.jsonlscripts/train.sh runs torchrun --nproc_per_node $NUM_GPUS src/train.py --config <config> (default NUM_GPUS=8). The paper uses a global batch size of 8 GPUs Γ 8 Γ 2 gradient-accumulation steps; when training on fewer GPUs, increase train.gradient_accumulation_steps to keep it unchanged. The final model is written to <output_dir>/final.
scripts/eval.sh embeds the queries, the original corpus, and the mixed corpus of every dataset with a single model load, then reports the Origin (per-dataset original corpus) and Mix (merged mixed-modality pool) results:
bash scripts/eval.sh <model_type> <checkpoint> [vdoc|natural] [extra options]Results (Recall / nDCG / MRR / MAP / Precision / Success @ {1, 5, 10}) are written to outputs/eval/<run>/:
<tag>_result.jsonβ Origin results per dataset,mix_result.jsonβ Mix results per dataset plus macro / micro averages,results_table.xlsxβ NDCG@10 (Γ100) in the layout of Table 1.
Supported model_types (see src/embedding_backends/):
model_type |
Model in the paper | Checkpoint |
|---|---|---|
trident_qwen3vl |
Trident-Qwen3VL-2B (ours) | <output_dir>/final (released checkpoint: coming soon) |
trident_jinaclip |
Trident-JinaCLIP (ours) | <output_dir>/final (released checkpoint: coming soon) |
qwen3vl_official |
Qwen3-VL-Embedding-2B / 8B | Qwen/Qwen3-VL-Embedding-2B, Qwen/Qwen3-VL-Embedding-8B |
jina_clip_v2 |
Jina-CLIP-v2 | jinaai/jina-clip-v2 |
jina_v5_omni |
Jina-embeddings-v5-omni-small | jinaai/jina-embeddings-v5-omni-small |
unime_phi35v |
UniME-Phi3.5-V-4.2B | DeepGlint-AI/UniME-Phi3.5-V-4.2B |
visrag_ret |
VisRAG-Ret | openbmb/VisRAG-Ret |
clip_vit_l14 |
CLIP-ViT-L/14 | openai/clip-vit-large-patch14 |
siglip2 |
SigLIP2-L/16-384 | google/siglip2-large-patch16-384 |
altclip |
AltCLIP | BAAI/AltCLIP |
Models without native joint textβimage encoding (CLIP-style dual towers, UniME, Jina-v5-omni) encode fused documents by averaging the normalized text and image embeddings with equal weights (--clip_fusion_alpha 0.5).
Benchmarks: ChartQA, DocVQA, InfoVQA, SlideVQA, ViDoSeek, Dude (NDCG@10, Origin and Mix).
# Trident-Qwen3VL-2B (LoRA adapter checkpoint + base model)
bash scripts/eval.sh trident_qwen3vl outputs/trident_qwen3vl_2b/final vdoc \
--base_model Qwen/Qwen3-VL-2B-Instruct
# Trident-JinaCLIP
bash scripts/eval.sh trident_jinaclip outputs/trident_jinaclip/final vdoc
# Baselines, e.g.
bash scripts/eval.sh qwen3vl_official Qwen/Qwen3-VL-Embedding-2B vdoc
bash scripts/eval.sh jina_clip_v2 jinaai/jina-clip-v2 vdoc"+ GR-CLIP" rows. GR-CLIP is a post-hoc mean-shift calibration for CLIP-based models. Compute the modality means once per model, then evaluate with GR_CLIP_MEANS set:
bash scripts/compute_gr_clip_means.sh jina_clip_v2 jinaai/jina-clip-v2
GR_CLIP_MEANS=outputs/gr_clip_calib/jina_clip_v2/gr_clip_means.npz \
bash scripts/eval.sh jina_clip_v2 jinaai/jina-clip-v2 vdocBenchmarks: Google-WIT, MSCOCO, VisualNews, OVEN (the Mix corpus follows MixBench). The models are trained on 170K samples from the MMEB training sets with the same configs; point data.train_path to the natural-image training file:
bash scripts/train.sh configs/trident_qwen3vl.yaml \
--set data.train_path=<natural-image train.jsonl> train.output_dir=outputs/trident_qwen3vl_2b_natural
bash scripts/eval.sh trident_qwen3vl outputs/trident_qwen3vl_2b_natural/final natural \
--base_model Qwen/Qwen3-VL-2B-InstructTraining (configs/*.yaml, any field can be overridden with --set key=value):
| Parameter | Meaning |
|---|---|
data.num_positives: 3 |
number of positive views per sample (text / image / fused). |
loss.modules.false_negative.mask_sibling_positives |
false: the positive views share one softmax denominator (Multi-Positive View InfoNCE, Eq. 4). true: each positive view is normalized independently against the negatives (no positive-view balance; ablation). |
loss.multi_positive_reduction |
mean (Eq. 4): average the log-probabilities of the positives; joint: only the total positive mass -log Ο. |
loss.temperature |
InfoNCE temperature Ο (0.02 for Qwen3-VL, 0.05 for jina-clip-v2). |
loss.symmetric |
add the document-to-query direction of InfoNCE. |
engine.cross_device: grad |
share in-batch negatives across GPUs with gradients (detach / off are cheaper alternatives). |
loss.modules.false_negative.* |
mask candidates that are actually positives of the query (same doc_id, positive_ids, group_id, or sibling view groups). |
loss.modules.modality_balance.weight |
weight of the explicit Modality Balance Loss (0 = log only, used by Trident; 200 in Appendix G). |
loss.modules.matryoshka |
Matryoshka training over several embedding dimensions. |
model.lora.*, model.freeze_vision |
LoRA rank / alpha and whether the vision tower is frozen. |
model.fusion_text_weight |
(jina-clip) weight of the text embedding when fusing text + image. |
model.image.max_pixels |
(Qwen3-VL) maximum image resolution; the main knob for GPU memory. |
Evaluation (scripts/run_all_datasets.sh, called by scripts/eval.sh):
| Option | Meaning |
|---|---|
--datasets A,B / --benchmark vdoc|natural |
datasets to evaluate. |
--no_mix / --only_mix |
run only the Origin or only the Mix evaluation. |
--gpus, --batch_size |
GPUs and per-GPU batch size for embedding. |
--base_model |
base model of a LoRA-adapter checkpoint (local path or HF repo id). |
--matryoshka_dims |
evaluate several truncated dimensions in one pass. |
--k_values |
cut-offs of the reported metrics (default 1,5,10). |
GR_CLIP_MEANS |
enable GR-CLIP calibration with the given means file (CLIP-based models only). |
Trident/
βββ configs/
β βββ base.yaml # shared defaults
β βββ trident_qwen3vl.yaml # Trident-Qwen3VL-2B
β βββ trident_jinaclip.yaml # Trident-JinaCLIP
β βββ ds_zero{2,3}.json # optional DeepSpeed configs (train.deepspeed=...)
βββ data/download_data.sh # data download script
βββ figs/ # figures used in this README
βββ scripts/
β βββ train.sh # training
β βββ eval.sh # Origin + Mix evaluation of one model
β βββ run_all_datasets.sh # evaluation pipeline used by eval.sh
β βββ compute_gr_clip_means.sh # GR-CLIP calibration means
βββ src/
β βββ train.py # training entry point
β βββ mmemb/ # training framework
β β βββ data/ # multi-view datasets, collator, false-negative identities
β β βββ models/ # Qwen3-VL-Embedding and jina-clip-v2 encoders
β β βββ losses/ # (Multi-Positive View) InfoNCE, false-negative masking, modality balance
β β βββ engine/ # HF-Trainer-based engine, cross-GPU negatives, monitoring
β βββ embedding_backends/ # inference backends of Trident and all baselines (+ GR-CLIP)
β βββ eval/ # multi-GPU embedding, Origin / Mix retrieval evaluation, result tables
β βββ tools/ # GR-CLIP mean computation
βββ requirements.txt
βββ LICENSE
This repository builds on transformers, PEFT, the official Qwen3-VL-Embedding model definition (Apache-2.0; adapted in src/mmemb/models/modeling_qwen3_vl_embedding.py), and jina-clip-v2. The GR-CLIP baseline follows Closing the Modality Gap for Mixed Modality Search. We thank the authors of OpenDocVQA, ChartQA, DocVQA, InfoVQA, SlideVQA, ViDoSeek, DUDE, MixBench, and MMEB for releasing their data.
This project is released under the MIT License. Models and datasets used in this project are subject to their own licenses.
If you find this work useful, please cite:
@article{trident2026chaos,
title = {Chaos in the Text: Revealing the Modality Preference in Mixed-Modality Retrievers},
author = {Sun, Yubo and Peng, Chunyi and Yan, Yukun and Liu, Zhenghao and Xu, Zhipeng and Mei, Sen and Xin, Linlin and Zeng, Zheni and Sun, Maosong},
journal = {arXiv preprint arXiv:2610.11816},
year = {2026},
url = {https://arxiv.org/abs/2610.11816}
}
