RallyMotionReasoner is a modular system for event-grounded tactical reasoning in tennis videos. It converts a rally into temporally aligned hit and bounce events, builds a tactical graph, retrieves evidence for a question, and optionally generates a grounded answer with a Qwen vision-language model.
The repository contains the model structure and integration code. It does not claim reproduced training results and does not include broadcast videos or pretrained weights.
The figure shows the intended joint event model. The current checkpointed runtime uses two independent experts with score-level fusion; the cross-attention and attribute heads shown in the figure are not active in that runtime.
The current event backend combines two separately checkpointed experts:
- TrajectoryExpert encodes ball position, velocity, acceleration, visibility, and temporal validity from TrackNet-style trajectories.
- VisualExpert encodes a full-frame DINO feature together with four spatial crops, preserving both court context and local player motion.
- Score-level fusion averages the experts' eventness and hit/bounce type scores.
- MotionRegionEventModel defines bidirectional cross-attention and optional attribute heads, but no checkpointed training or inference path uses it yet.
The runtime loads trajectory_expert.pt and visual_expert.pt from one expert directory. It decodes hit and bounce events with temporal suppression. It does not currently predict hitter or technique attributes. Bounce events remain separate graph cues; they are not converted into player stroke nodes.
Hit events become stroke nodes. Bounce events provide timing context such as the gap before or after a stroke. RGR receives:
- 800-dimensional frame descriptors that include 24-dimensional motion statistics;
- shot timing and event attributes when available;
- temporal relations between neighboring strokes;
- same-player relations where the hitter is known (the current event runtime does not supply hitter labels);
- bounce-aware structural tokens.
The graph loader and event exporter use the same structured feature payload, so offline RGR data and online inference share the frame alignment contract.
The evidence router scores candidate strokes and key actions for a tactical question. Key-action probability is computed from the conditional key score and evidence score. The generator receives validated candidate identifiers, their evidence frames, sparse global context, and predicted event information. It must return the canonical answer fields defined in src/rallymotionreasoner/schema.py.
- Python 3.10 or newer
- PyTorch 2.1 or newer
- OpenCV,
ffprobe, and DINOv3 dependencies for raw-video event extraction - CUDA is recommended for model inference and required by most large-model training setups
Install the package and optional training dependencies:
git clone https://github.com/ZSHYC/RallyMotionReasoner.git
cd RallyMotionReasoner
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[train,generation,video]'Copy the path template and fill in local data and model locations when using the data preparation or training scripts:
cp configs/paths.example.yaml configs/paths.local.yamlA ball trajectory can be a TrackNet JSON payload or a CSV file. Trajectory frames must be ordered and finite, with coordinates inside the track's declared source dimensions. Raw-video inputs also check those dimensions against the decoded video. Raw videos must have constant frame rate; decoded frames are stored temporarily as lossless PNGs. A frame directory requires an explicit --fps and may contain resized copies of the source frames. The event expert directory must contain:
region-experts/
├── trajectory_expert.pt
└── visual_expert.pt
The video, trajectory, DINOv3 repository/weights, RGR checkpoint, and Qwen model are supplied independently.
Show all available commands:
PYTHONPATH=src python -m rallymotionreasoner.cli --helpPYTHONPATH=src python -m rallymotionreasoner.cli prepare-data \
--config configs/rallymotionreasoner.yamlThis validates configured splits and builds graph/QA artifacts. It does not train a model.
rallymotionreasoner predict-events \
--video /path/to/rally.mp4 \
--ball-track /path/to/trajectory.json \
--checkpoint /path/to/region-experts \
--dinov3-repo /path/to/dinov3 \
--dinov3-weights /path/to/dinov3.pth \
--output outputs/events.jsonUse PYTHONPATH=src python -m rallymotionreasoner.cli predict-events when the editable package is not installed. The output contains decoded events, frame scores, frame features, and feature provenance.
rallymotionreasoner export-events \
--checkpoint /path/to/region-experts \
--split train \
--experiment-config configs/rallymotionreasoner.yaml \
--dinov3-repo /path/to/dinov3 \
--dinov3-weights /path/to/dinov3.pth \
--output outputs/events_trainThe exporter reads configured rally videos and TrackNet files, writes predicted event graphs, and stores hit-only frame features using the shared RGR payload format. Run it separately for train, val, and test when those splits are configured.
rallymotionreasoner predict \
--video /path/to/rally.mp4 \
--ball-track /path/to/trajectory.json \
--event-checkpoint /path/to/region-experts \
--rgr-checkpoint /path/to/rgr.pt \
--qwen-model /path/to/Qwen3-VL-8B-Instruct \
--question "How did the player create the winning opportunity?" \
--dinov3-repo /path/to/dinov3 \
--dinov3-weights /path/to/dinov3.pth \
--output outputs/answer.jsonThe full pipeline requires a Qwen model, supplied through --qwen-model or qwen3_vl_model in the paths configuration. The Qwen adapter is optional. Use predict-events for event-only output.
Training is separate from event inference and is not required for reading the model structure:
PYTHONPATH=src python scripts/train_graph_reasoner.py \
--experiment-config configs/rallymotionreasoner.yaml
PYTHONPATH=src torchrun --nproc-per-node=8 scripts/train_qwen_lora.py \
--model /path/to/Qwen3-VL-8B-Instruct \
--train-data /path/to/train.jsonl \
--val-data /path/to/val.jsonl \
--data-report /path/to/qwen_data_report.json \
--output-dir runs/qwen_lora \
--report outputs/qwen_lora.jsonTraining requires locally prepared data and model weights. Qwen SFT rows contain one frame-list video, a canonical predicted-candidate prompt without a system message, original video frame indices and FPS, and candidate centers included in those frames. Assistant evidence IDs and frames must match prompt candidates. The data report must cover both splits and declare the 32-frame sampling policy. Qwen adapters carry that prompt, sampling, and source-timing contract in their manifest; resumed training also checks the data paths and effective row counts.
assets/architecture.png system architecture diagram
configs/ YAML experiment and path configuration
scripts/train_graph_reasoner.py RGR training entry point
scripts/train_qwen_lora.py Qwen LoRA training entry point
src/rallymotionreasoner/event_detection/ event model, features, graph, decoder, runtime
src/rallymotionreasoner/features/ trajectory input and normalization
src/rallymotionreasoner/graph_reasoning/ RGR, motion adapter, evidence routing
src/rallymotionreasoner/generation/ structured generation and manifest handling
src/rallymotionreasoner/pipeline.py end-to-end event → RGR → generation flow
tests/ lightweight contract and shape tests
docs/algorithm.md algorithm and interface notes
Run the lightweight checks used for code changes:
ruff check src scripts tests
PYTHONPATH=src pytest -qThe tests validate tensor shapes, masked fusion behavior, event graph semantics, feature alignment, RGR contracts, configuration validation, and structured output fields. They do not run training or real video inference.
See LICENSE for the applicable license and attribution requirements.
