Skip to content

feat: Whitepaper benchmarks with CI exclusion and confidence intervals - #2

Merged
Ricoledan merged 1 commit into
mainfrom
feat/whitepaper-benchmarks
Feb 11, 2026
Merged

Ricoledan merged 1 commit into
mainfrom
feat/whitepaper-benchmarks

Conversation

@Ricoledan

Copy link
Copy Markdown
Member

Summary

  • Whitepaper benchmark harness (tests/performance/test_whitepaper_benchmarks.py): 21+ benchmarks producing empirical evidence for three paper sections — performance overhead, detection effectiveness, and breach prevention against 2025 MCP attack scenarios
  • Multi-backend inference distribution (v0.3.0): OpenAICompatibleClient base class with vLLM, SGLang, and llama.cpp backend support; create_llm_client() factory for config-driven selection
  • 95% confidence intervals: _compute_stats() now returns ci95_lower/ci95_upper/stdev/n via t-distribution; _aggregate_multi_run_stats() for cross-run reproducibility
  • CI exclusion: Benchmarks marked @pytest.mark.benchmark are skipped in CI ("not docker and not benchmark")

Test plan

  • pytest -m "not docker and not benchmark" --co -q collects 0 whitepaper tests
  • pytest tests/performance/test_whitepaper_benchmarks.py --co -q collects 23 tests
  • Single benchmark run produces JSON with ci95_lower, ci95_upper, stdev, n keys
  • ruff check and all pre-commit hooks pass
  • CI passes with benchmark exclusion

🤖 Generated with Claude Code

Update _compute_stats() to return ci95_lower/ci95_upper/stdev/n using
t-distribution CIs for paper-grade statistical rigor. Add
_aggregate_multi_run_stats() for cross-run reproducibility evidence.
@Ricoledan
Ricoledan merged commit da05fb7 into main Feb 11, 2026
7 checks passed
@Ricoledan
Ricoledan deleted the feat/whitepaper-benchmarks branch February 11, 2026 03:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant