Skip to content

Latest commit

 

History

41 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

RallyMotionReasoner

RallyMotionReasoner is a modular system for event-grounded tactical reasoning in tennis videos. It converts a rally into temporally aligned hit and bounce events, builds a tactical graph, retrieves evidence for a question, and optionally generates a grounded answer with a Qwen vision-language model.

The repository contains the model structure and integration code. It does not claim reproduced training results and does not include broadcast videos or pretrained weights.

What the system does

RallyMotionReasoner architecture: event detection, tactical graph, RGR reasoning, and grounded generation

The figure shows the intended joint event model. The current checkpointed runtime uses two independent experts with score-level fusion; the cross-attention and attribute heads shown in the figure are not active in that runtime.

Event detection

The current event backend combines two separately checkpointed experts:

  • TrajectoryExpert encodes ball position, velocity, acceleration, visibility, and temporal validity from TrackNet-style trajectories.
  • VisualExpert encodes a full-frame DINO feature together with four spatial crops, preserving both court context and local player motion.
  • Score-level fusion averages the experts' eventness and hit/bounce type scores.
  • MotionRegionEventModel defines bidirectional cross-attention and optional attribute heads, but no checkpointed training or inference path uses it yet.

The runtime loads trajectory_expert.pt and visual_expert.pt from one expert directory. It decodes hit and bounce events with temporal suppression. It does not currently predict hitter or technique attributes. Bounce events remain separate graph cues; they are not converted into player stroke nodes.

Graph and RGR

Hit events become stroke nodes. Bounce events provide timing context such as the gap before or after a stroke. RGR receives:

  • 800-dimensional frame descriptors that include 24-dimensional motion statistics;
  • shot timing and event attributes when available;
  • temporal relations between neighboring strokes;
  • same-player relations where the hitter is known (the current event runtime does not supply hitter labels);
  • bounce-aware structural tokens.

The graph loader and event exporter use the same structured feature payload, so offline RGR data and online inference share the frame alignment contract.

Evidence-grounded generation

The evidence router scores candidate strokes and key actions for a tactical question. Key-action probability is computed from the conditional key score and evidence score. The generator receives validated candidate identifiers, their evidence frames, sparse global context, and predicted event information. It must return the canonical answer fields defined in src/rallymotionreasoner/schema.py.

Requirements

  • Python 3.10 or newer
  • PyTorch 2.1 or newer
  • OpenCV, ffprobe, and DINOv3 dependencies for raw-video event extraction
  • CUDA is recommended for model inference and required by most large-model training setups

Install the package and optional training dependencies:

git clone https://github.com/ZSHYC/RallyMotionReasoner.git
cd RallyMotionReasoner
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[train,generation,video]'

Copy the path template and fill in local data and model locations when using the data preparation or training scripts:

cp configs/paths.example.yaml configs/paths.local.yaml

Inputs

A ball trajectory can be a TrackNet JSON payload or a CSV file. Trajectory frames must be ordered and finite, with coordinates inside the track's declared source dimensions. Raw-video inputs also check those dimensions against the decoded video. Raw videos must have constant frame rate; decoded frames are stored temporarily as lossless PNGs. A frame directory requires an explicit --fps and may contain resized copies of the source frames. The event expert directory must contain:

region-experts/
├── trajectory_expert.pt
└── visual_expert.pt

The video, trajectory, DINOv3 repository/weights, RGR checkpoint, and Qwen model are supplied independently.

Commands

Show all available commands:

PYTHONPATH=src python -m rallymotionreasoner.cli --help

Prepare graph and QA data

PYTHONPATH=src python -m rallymotionreasoner.cli prepare-data \
  --config configs/rallymotionreasoner.yaml

This validates configured splits and builds graph/QA artifacts. It does not train a model.

Run event detection

rallymotionreasoner predict-events \
  --video /path/to/rally.mp4 \
  --ball-track /path/to/trajectory.json \
  --checkpoint /path/to/region-experts \
  --dinov3-repo /path/to/dinov3 \
  --dinov3-weights /path/to/dinov3.pth \
  --output outputs/events.json

Use PYTHONPATH=src python -m rallymotionreasoner.cli predict-events when the editable package is not installed. The output contains decoded events, frame scores, frame features, and feature provenance.

Export events for RGR

rallymotionreasoner export-events \
  --checkpoint /path/to/region-experts \
  --split train \
  --experiment-config configs/rallymotionreasoner.yaml \
  --dinov3-repo /path/to/dinov3 \
  --dinov3-weights /path/to/dinov3.pth \
  --output outputs/events_train

The exporter reads configured rally videos and TrackNet files, writes predicted event graphs, and stores hit-only frame features using the shared RGR payload format. Run it separately for train, val, and test when those splits are configured.

Run the complete pipeline

rallymotionreasoner predict \
  --video /path/to/rally.mp4 \
  --ball-track /path/to/trajectory.json \
  --event-checkpoint /path/to/region-experts \
  --rgr-checkpoint /path/to/rgr.pt \
  --qwen-model /path/to/Qwen3-VL-8B-Instruct \
  --question "How did the player create the winning opportunity?" \
  --dinov3-repo /path/to/dinov3 \
  --dinov3-weights /path/to/dinov3.pth \
  --output outputs/answer.json

The full pipeline requires a Qwen model, supplied through --qwen-model or qwen3_vl_model in the paths configuration. The Qwen adapter is optional. Use predict-events for event-only output.

Train downstream components

Training is separate from event inference and is not required for reading the model structure:

PYTHONPATH=src python scripts/train_graph_reasoner.py \
  --experiment-config configs/rallymotionreasoner.yaml

PYTHONPATH=src torchrun --nproc-per-node=8 scripts/train_qwen_lora.py \
  --model /path/to/Qwen3-VL-8B-Instruct \
  --train-data /path/to/train.jsonl \
  --val-data /path/to/val.jsonl \
  --data-report /path/to/qwen_data_report.json \
  --output-dir runs/qwen_lora \
  --report outputs/qwen_lora.json

Training requires locally prepared data and model weights. Qwen SFT rows contain one frame-list video, a canonical predicted-candidate prompt without a system message, original video frame indices and FPS, and candidate centers included in those frames. Assistant evidence IDs and frames must match prompt candidates. The data report must cover both splits and declare the 32-frame sampling policy. Qwen adapters carry that prompt, sampling, and source-timing contract in their manifest; resumed training also checks the data paths and effective row counts.

Repository layout

assets/architecture.png         system architecture diagram
configs/                         YAML experiment and path configuration
scripts/train_graph_reasoner.py            RGR training entry point
scripts/train_qwen_lora.py       Qwen LoRA training entry point
src/rallymotionreasoner/event_detection/     event model, features, graph, decoder, runtime
src/rallymotionreasoner/features/           trajectory input and normalization
src/rallymotionreasoner/graph_reasoning/ RGR, motion adapter, evidence routing
src/rallymotionreasoner/generation/        structured generation and manifest handling
src/rallymotionreasoner/pipeline.py        end-to-end event → RGR → generation flow
tests/                           lightweight contract and shape tests
docs/algorithm.md                algorithm and interface notes

Validation

Run the lightweight checks used for code changes:

ruff check src scripts tests
PYTHONPATH=src pytest -q

The tests validate tensor shapes, masked fusion behavior, event graph semantics, feature alignment, RGR contracts, configuration validation, and structured output fields. They do not run training or real video inference.

License

See LICENSE for the applicable license and attribution requirements.

About

Rally-level motion reasoning and evidence-grounded tactical understanding for tennis videos

Resources

Stars

52 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages