Conversation
Turn the trace artifacts a run leaves behind into TraceLens operator, kernel, roofline, and collective reports, and add on-demand torch.profiler capture so PyTorch workloads can produce those traces without editing the model script. Analysis runs in either of two places. On the host, `madengine report tracelens` keeps TraceLens' pinned protobuf and xprof out of workload images entirely. In-container, the `tracelens` tool installs into an isolated virtualenv for the same reason. SLURM and Kubernetes collection now gather torch_profiler_output/ and tracelens_output/ per node. The `scripts/` and `*.json` ignore rules are anchored to the repo root: unanchored they also matched packaged source, which silently excluded the five runtime scripts these tool definitions depend on. Co-authored-by: Cursor <cursoragent@cursor.com>
There was a problem hiding this comment.
Pull request overview
Integrates TraceLens-based trace analysis into madengine, providing both host-side reporting (madengine report tracelens / tracelens-compare) and in-container tooling to capture PyTorch Kineto traces on-demand (dynolog) and generate TraceLens reports. This fits into the reporting + tools pipeline by turning collected profiler artifacts into actionable operator/kernel/roofline/collective outputs and ensuring distributed collectors gather the needed directories.
Changes:
- Add a stdlib-only TraceLens analyzer script plus a host-side wrapper module and new report CLI commands (
tracelens,tracelens-compare). - Add in-container tools (
torch_profiler_dynolog,tracelensand mode-specific variants) with supporting pre/post scripts and TraceLens-venv isolation. - Extend SLURM/Kubernetes artifact collection, add examples/docs, and add unit/integration/e2e coverage.
Reviewed changes
Copilot reviewed 24 out of 25 changed files in this pull request and generated 3 comments.
Show a summary per file
| File | Description |
|---|---|
| tests/unit/test_tracelens_report.py | Unit tests for host-side TraceLens wrapper functions and report CLI behavior. |
| tests/unit/test_tracelens_analyze.py | Unit tests for the standalone analyzer script (discovery, command construction, summaries). |
| tests/integration/test_tracelens_tools_config.py | Integration tests validating tools.json wiring and tool stacking behavior. |
| tests/e2e/test_tracelens_workflows.py | E2E coverage for host reporting and (GPU) container tool workflows. |
| src/madengine/scripts/common/tools/tracelens_analyze.py | Standalone (stdlib-only) analyzer that discovers traces and runs TraceLens entry points out-of-process. |
| src/madengine/scripts/common/tools/dynolog_trigger.sh | Background trigger script to request torch.profiler traces via dynolog. |
| src/madengine/scripts/common/tools.json | Registers torch_profiler_dynolog and TraceLens tool variants and their scripts/env. |
| src/madengine/scripts/common/pre_scripts/trace.sh | Adds installers for dynolog and TraceLens (isolated venv + optional traceconv staging). |
| src/madengine/scripts/common/pre_scripts/dynolog_start.sh | Starts dynolog daemon and arms the trigger alongside the workload. |
| src/madengine/scripts/common/post_scripts/tracelens.sh | Runs the analyzer post-workload to generate TraceLens reports without failing the run. |
| src/madengine/scripts/common/post_scripts/trace.sh | Collects new output directories (torch_profiler_output, tracelens_output) as artifacts. |
| src/madengine/scripts/common/post_scripts/dynolog_stop.sh | Stops dynolog/trigger and reports captured trace count + tail logs. |
| src/madengine/reporting/tracelens_report.py | Host-side driver for running the analyzer and surfacing install guidance/errors. |
| src/madengine/deployment/templates/slurm/job.sh.j2 | Collects additional per-node tool output directories on successful SLURM runs. |
| src/madengine/deployment/templates/kubernetes/job.yaml.j2 | Copies additional per-pod tool output directories into results for K8s runs. |
| src/madengine/deployment/k8s_scripts.py | Ensures helper scripts referenced by pre/post scripts under scripts/common/tools are bundled for K8s. |
| src/madengine/deployment/k8s_results.py | Collects .pftrace and additional known output directories via kubectl cp. |
| src/madengine/cli/commands/report.py | Adds report tracelens and report tracelens-compare commands with Rich output. |
| pyproject.toml | Adds optional dependency extra madengine[tracelens] pinned to a TraceLens git ref. |
| examples/profiling-configs/tracelens_rocprofv3.json | Example config for rocprofv3_lightweight + TraceLens analysis. |
| examples/profiling-configs/torch_profiler_tracelens.json | Example config for dynolog capture + TraceLens PyTorch analysis. |
| examples/profiling-configs/README.md | Documents TraceLens/dynolog example configs and produced artifacts. |
| docs/profiling.md | Adds documentation for dynolog-based capture and TraceLens host/container analysis workflows. |
| docs/cli-reference.md | Adds CLI reference entries for report tracelens and report tracelens-compare. |
| .gitignore | Anchors scripts/ and *.json ignore rules to repo root to avoid ignoring packaged runtime scripts. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| if [ "$(id -u)" -eq 0 ]; then | ||
| dpkg -i "$_dynolog_tmp" || apt-get install -f -y -qq || true | ||
| elif command -v sudo >/dev/null 2>&1; then | ||
| sudo dpkg -i "$_dynolog_tmp" || sudo apt-get install -f -y -qq || true |
| return _run_analyzer(args, summary_path) | ||
|
|
||
|
|
||
| def discover_traces(root: str = ".", output_dir: str = "tracelens_output") -> Dict[str, object]: |
| with open(abs_script_path, "r") as f: | ||
| script_content = f.read() | ||
| tool_refs = re.findall(r'(?:\.\./)?scripts/common/tools/[\w_]+\.py', script_content) | ||
| tool_refs = re.findall(r'(?:\.\./)?scripts/common/tools/[\w_]+\.(?:py|sh)', script_content) | ||
| for tool_ref in tool_refs: |
| return _run_analyzer(args, summary_path) | ||
|
|
||
|
|
||
| def discover_traces(root: str = ".", output_dir: str = "tracelens_output") -> Dict[str, object]: |
There was a problem hiding this comment.
Black would reformat this signature (confirmed with black --check):
def discover_traces(
root: str = ".", output_dir: str = "tracelens_output"
) -> Dict[str, object]:
Repo convention is Black formatting (88-char line length per CLAUDE.md); this will fail pre-commit run --all-files/CI lint. Worth running black src/ tests/ before merge.
| # harmless because we run the daemon directly. Tolerate a non-zero dpkg exit | ||
| # and verify by checking for the binaries instead. | ||
| if [ "$(id -u)" -eq 0 ]; then | ||
| dpkg -i "$_dynolog_tmp" || apt-get install -f -y -qq || true |
There was a problem hiding this comment.
This dpkg fallback runs apt-get install -f -y -qq without a preceding apt-get update, so on images with stale/empty package lists this can't actually resolve dynolog's missing deps. Not silent -- the script does verify afterward with command -v dynolog/dyno and exits with a clear error if install failed -- but it's an avoidable extra failure mode on minimal base images. Suggest adding apt-get update right before this fallback (same for the sudo branch on line 190).
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Critical credential exposure risks and blocking profiling, analysis, collection, and CLI failures remain unresolved.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 3
Open (9)
Prevent Dynolog URL credentials from xtrace output · New Prevent TraceLens credential leakage through xtrace · New Remove unsupported Dynolog fail-on-no-process flag · New Treat analyzer subprocess failures as report failures · New Match Dynolog libkineto trace filenames · New Restrict collective analysis to discovered PyTorch traces · New Dynolog install path runsapt-get install -fwithout anapt-get updatefirst. On minimal… Thisre.findall(...)line exceeds Black’s configured 88-character limit, which can cause… This function definition exceeds the configured Black line length (88 chars), which can cause…
| _dynolog_deb="${DYNOLOG_DEB_URL:-$_DYNOLOG_PINNED_DEB}" | ||
| _dynolog_tmp="/tmp/dynolog.deb" | ||
|
|
||
| if command -v curl >/dev/null 2>&1; then | ||
| curl -fsSL -o "$_dynolog_tmp" "$_dynolog_deb" |
| _tl_venv="${TRACELENS_VENV:-/opt/madengine-tracelens-venv}" | ||
| _TRACELENS_PINNED_REF='6f9bcdbf6cc9911eb650de57b345917ea4d31a17' | ||
| _tl_ref="${TRACELENS_GIT_REF:-$_TRACELENS_PINNED_REF}" | ||
| _tl_spec="${TRACELENS_PIP_SPEC:-git+https://github.com/AMD-AGI/TraceLens.git@${_tl_ref}}" |
| if dyno --port "$PORT" gputrace \ | ||
| --job-id "$JOB_ID" \ | ||
| --log-file "$LOG_FILE" \ | ||
| --process-limit "$PROCESS_LIMIT" \ | ||
| --fail-on-no-process \ | ||
| "${OPTS[@]}"; then | ||
| echo "[dynolog-trigger] trace request accepted on attempt ${attempt}" | ||
| echo "accepted" > "$RESULT_FILE" | ||
| exit 0 | ||
| fi |
| succeeded = int(summary.get("succeeded", 0)) | ||
| failed = int(summary.get("failed", 0)) | ||
| skipped = int(summary.get("skipped", 0)) |
| renames per rank), torch.profiler defaults (``..._rank0_...``), and the | ||
| ``rank[N]`` form used by ``tensorboard_trace_handler``. | ||
| """ | ||
| return r"rank[\[\-_/]?(?P<rank>\d+)" |
| "--trace_glob", | ||
| os.path.join(root, "**", "*.json*"), |



Turn the trace artifacts a run leaves behind into TraceLens operator, kernel, roofline, and collective reports, and add on-demand torch.profiler capture so PyTorch workloads can produce those traces without editing the model script.
Analysis runs in either of two places. On the host,
madengine report tracelenskeeps TraceLens' pinned protobuf and xprof out of workload images entirely. In-container, thetracelenstool installs into an isolated virtualenv for the same reason. SLURM and Kubernetes collection now gather torch_profiler_output/ and tracelens_output/ per node.The
scripts/and*.jsonignore rules are anchored to the repo root: unanchored they also matched packaged source, which silently excluded the five runtime scripts these tool definitions depend on.