diff --git a/CHANGELOG.md b/CHANGELOG.md index e69de29..5c459bb 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -0,0 +1,23 @@ +# DevOS AI Changelog + +## v3.0.0 - 2026-04-12 + +### πŸš€ Features +- Added a new `core/v3` module with weighted context ranking and modern prompt templates. +- Added advanced explain controls to the CLI: `--max-files`, `--max-chars`, `--include-tests`, and `--json`. +- Added explain output metadata showing the selected files used as context. + +### πŸ”„ Replacements +- Directly replaced the old simplistic context-building flow with the new v3 engine pipeline. +- Replaced ad-hoc CLI argument handling with `argparse` for more reliable command parsing. + +### πŸ›  Improvements +- Expanded parser extension support and safer file reading behavior. +- Added release notes under `VERSION/3.0.0/CHANGELOG.md`. + +## v2.0.0 +- CLI-based AI code analyzer +- Code explanation +- Code search +- Debug assistant +- Multi-LLM support diff --git a/README.md b/README.md index 6307a62..931afdd 100644 --- a/README.md +++ b/README.md @@ -1,228 +1,73 @@ # πŸš€ DevOS AI -CLI-based AI tool for analyzing and explaining codebases using structured context extraction and LLM reasoning. +DevOS AI is a developer-first assistant for understanding codebases faster using structured context extraction and LLM reasoning. ---- +## βœ… Release Status -## ✨ Introduction +**Current release: v3.0.0** -DevOS AI is a developer-first tool designed to help engineers quickly understand unfamiliar codebases. +v3 introduces a direct replacement of the legacy context engine with a smarter ranking pipeline, advanced CLI controls, and better output structure for explain workflows. -Instead of manually reading hundreds of files, DevOS intelligently extracts relevant parts of a repository, builds structured context, and uses Large Language Models (LLMs) to generate clear, high-level explanations of architecture and functionality. +## πŸ”₯ What’s New in v3 ---- - -## πŸ“š Table of Contents - -- Overview -- Features -- How It Works -- Architecture -- Project Structure -- Installation -- Usage -- Example Output -- Demo -- Roadmap -- Configuration -- Dependencies -- Contributing -- Troubleshooting -- License -- Author - ---- - -## πŸ” Overview - -DevOS AI transforms codebases into understandable insights by combining: - -- Smart file selection -- Context-aware prompt generation -- LLM-powered reasoning - -πŸ‘‰ It acts like a **codebase explainer engine**, helping developers onboard faster and debug smarter. - ---- - -## πŸ”₯ Features - -- πŸ” Analyze any codebase using AI -- 🧠 Structured context extraction -- πŸ— Architecture-level explanations -- βš™οΈ Modular design (CLI + Core + LLM) -- 🌐 Multi-LLM support (OpenRouter + Google) -- πŸ’» CLI-first developer experience - ---- - -## 🧠 How It Works - -### Pipeline - - -CLI β†’ Context Builder β†’ Prompt Engine β†’ LLM β†’ Output - - -### Step-by-Step - -1. Extract important files from the repository -2. Rank and filter relevant code -3. Build structured context -4. Generate optimized prompts -5. Send to LLM for reasoning -6. Output structured explanation - ---- - -## πŸ— Architecture - -| Module | Description | -|--------|------------| -| `cli/` | Command-line interface | -| `core/` | Context building + prompt generation | -| `llm/` | LLM provider integration | - ---- - -## πŸ“‚ Project Structure +- Advanced context ranking with weighted file scoring. +- Direct replacement of the old prompt builder with a stricter β€œfacts vs inference” response format. +- New `core/v3/` module for maintainable architecture. +- New CLI controls: `--max-files`, `--max-chars`, `--include-tests`, and `--json`. +- Improved explain command output with selected file visibility. +## πŸ“¦ Project Structure +```text devos-ai/ β”œβ”€β”€ cli/ β”œβ”€β”€ core/ +β”‚ └── v3/ β”œβ”€β”€ llm/ β”œβ”€β”€ agents/ β”œβ”€β”€ apps/ β”œβ”€β”€ frontend/ β”œβ”€β”€ infrastructure/ β”œβ”€β”€ docs/ -β”œβ”€β”€ scripts/ - - ---- +└── scripts/ +``` ## βš™οΈ Installation -### Prerequisites - -- Python 3.9+ -- API key (OpenRouter or Google Gemini) - -### Setup - +```bash git clone https://github.com/dsk-dev-ai/devos-ai.git cd devos-ai pip install -r requirements.txt +``` ## ⚑ Usage -πŸ” Explain Codebase -python3 -m cli.main explain . +### Explain a repository -πŸ”Ž Search Code -python3 -m cli.main search . "engine" - -🐞 Debug Error -python3 -m cli.main debug error.log - -🧠 Choose Model -python3 -m cli.main explain . --model openrouter -python3 -m cli.main explain . --model google +```bash python3 -m cli.main explain . --model auto +python3 -m cli.main explain . --max-files 12 --max-chars 2000 +python3 -m cli.main explain . --include-tests --json +``` -## 🧠 Example Output - -1. Purpose -CLI-based AI tool for analyzing codebases... - -2. Architecture Overview -CLI β†’ Core β†’ LLM β†’ Output - -3. Key Components -- CLI -- Core Engine -- LLM Provider - -### Example: -At (assets/images/) & (assets/videos/) - -## πŸš€ Roadmap - -βœ… V2 (Current) -CLI-based code analyzer -Context ranking system -Multi-LLM support -Search + Debug commands - -πŸ”œ V3 (Upcoming) -πŸ” Advanced file-level search -🐞 Smart stack trace debugging -⚑ RAG (vector search) -πŸ“¦ Global CLI install (devos) -🌐 Web dashboard - -## βš™οΈ Configuration - -Create .env file: - -OPENROUTER_API_KEY=your_key -GOOGLE_API_KEY=your_key +### Search code -## πŸ“¦ Dependencies - -Python standard library -requests -python-dotenv - -## 🀝 Contributing - -Contributions are welcome! +```bash +python3 -m cli.main search . "engine" +``` -You can help with: +### Debug an error file -Improving context extraction -Enhancing LLM prompts -Adding new CLI features -UI/UX improvements +```bash +python3 -m cli.main debug error.log +``` -## πŸ›  Troubleshooting +## πŸ›£οΈ Roadmap -1. LLM not working -Check API keys in .env -Verify internet connection -2. No output -Ensure valid project path -Check file permissions -3. Timeout / Slow response -Reduce context size -Check API limits +- βœ… v3: advanced context engine and richer CLI controls. +- πŸ”œ v3.x: smarter semantic retrieval and local embedding cache. +- πŸ”œ v4: interactive web dashboard with collaborative review mode. ## πŸ“œ License -This project is licensed under the MIT License. - -## πŸ‘¨β€πŸ’» Author - -Darshan Kachare -AI Developer β€’ Open Source Contributor - -GitHub: https://github.com/dsk-dev-ai - -## ⭐ Support - -If you find this useful: - -⭐ Star the repository -πŸ› Report issues -🀝 Contribute -πŸš€ Project Status - -## 🚧 Project Status - -- Experimental CLI-based AI tool -- Actively under development (v2) -- Focused on improving context extraction and LLM reasoning -- Suitable for learning and experimentation - ---- +MIT License. diff --git a/VERSION/3.0.0/CHANGELOG.md b/VERSION/3.0.0/CHANGELOG.md new file mode 100644 index 0000000..aa4c38f --- /dev/null +++ b/VERSION/3.0.0/CHANGELOG.md @@ -0,0 +1,15 @@ +# DevOS AI v3.0.0 + +## Highlights +- New `core/v3` architecture for advanced context construction. +- Direct replacement of the previous engine internals with strongly typed v3 builders. +- Explain command now supports: + - `--max-files` + - `--max-chars` + - `--include-tests` + - `--json` + +## Technical Notes +- Legacy `core/engine.py` functions are kept as compatibility wrappers and now route to v3 internals. +- File scoring has been upgraded to keyword + file-type weighting. +- Prompt output now includes risks and gaps for practical code review. diff --git a/cli/main.py b/cli/main.py index 8e94ab0..e7e5747 100755 --- a/cli/main.py +++ b/cli/main.py @@ -1,51 +1,70 @@ -import sys -from core.engine import build_context, build_prompt -from llm.provider import ask_llm +import argparse +import json + from agents.debug_agent import debug_error from agents.search_agent import search_code +from core.engine import build_context_with_files, build_prompt +from llm.provider import ask_llm -def explain(path, model): - context = build_context(path) - prompt = build_prompt(context, "Explain this codebase") - return ask_llm(prompt, model) - +def explain(path: str, model: str, max_files: int, max_chars: int, include_tests: bool) -> tuple[str, list[str]]: + context, selected_files = build_context_with_files( + path, + max_files=max_files, + max_chars=max_chars, + include_tests=include_tests, + ) + prompt = build_prompt(context, "Explain this codebase", selected_files) + return ask_llm(prompt, model), selected_files -def main(): - if len(sys.argv) < 3: - print(""" -Usage: - explain [--model openrouter|google|auto] - search "" [--model openrouter|google|auto] - debug [--model openrouter|google|auto] -""") - return - command = sys.argv[1] - model = "auto" +def build_parser() -> argparse.ArgumentParser: + parser = argparse.ArgumentParser(description="DevOS AI CLI v3") + parser.add_argument("command", choices=["explain", "search", "debug"], help="Command to run") + parser.add_argument("target", help="Repository path for explain/search, file path for debug") + parser.add_argument("query", nargs="?", default=None, help="Search query") + parser.add_argument("--model", default="auto", choices=["openrouter", "google", "auto"]) + parser.add_argument("--max-files", type=int, default=8) + parser.add_argument("--max-chars", type=int, default=1400) + parser.add_argument("--include-tests", action="store_true") + parser.add_argument("--json", action="store_true", dest="as_json") + return parser - if "--model" in sys.argv: - idx = sys.argv.index("--model") - if idx + 1 < len(sys.argv): - model = sys.argv[idx + 1] - print("πŸ” Processing...\n") +def main() -> None: + args = build_parser().parse_args() - if command == "explain": - result = explain(sys.argv[2], model) + if args.command == "explain": + result, selected_files = explain( + args.target, + args.model, + args.max_files, + args.max_chars, + args.include_tests, + ) - elif command == "search": - result = search_code(sys.argv[2], sys.argv[3], model) + if args.as_json: + print(json.dumps({"selected_files": selected_files, "result": result}, indent=2)) + return - elif command == "debug": - result = debug_error(sys.argv[2], model) + elif args.command == "search": + if not args.query: + raise SystemExit("search command requires a query") + result = search_code(args.target, args.query, args.model) + selected_files = None else: - result = "❌ Unknown command" + result = debug_error(args.target, args.model) + selected_files = None print("πŸ’‘ Result:\n") print(result) + if selected_files: + print("\nπŸ“ Selected files:") + for file_path in selected_files: + print(f"- {file_path}") + if __name__ == "__main__": main() diff --git a/core/engine.py b/core/engine.py index d1e29c1..71c6fec 100644 --- a/core/engine.py +++ b/core/engine.py @@ -1,64 +1,38 @@ -from core.parser import get_code_files, read_files - -PRIORITY = ["main", "app", "engine", "cli", "api"] - - -def rank_files(files): - scored = [] - - for f in files: - score = 0 - name = f.lower() - - for p in PRIORITY: - if p in name: - score += 5 - - if "core" in name or "llm" in name: - score += 3 - - scored.append((score, f)) - - scored.sort(reverse=True) - return [f for _, f in scored] - - -def build_context(repo_path): - files = get_code_files(repo_path) - ranked = rank_files(files) - - top_files = ranked[:5] - return read_files(top_files) - - -def build_prompt(context, question): - return f""" -You are a senior software engineer analyzing a real codebase. - -STRICT RULES: -- Use ONLY the given code context -- Do NOT assume missing features -- Do NOT mention technologies not present in context - -CONTEXT: -{context} - -TASK: -{question} - -Respond in this format: - -1. Purpose -- What this system actually does (based only on code) - -2. Architecture -- Real structure (modules, flow) - -3. Key Components -- Explain actual files/modules found - -4. Execution Flow -- Step-by-step how system works - -Keep answer precise, technical, and grounded in code. -""" +from __future__ import annotations + +from core.v3 import BuildOptions +from core.v3 import build_context as build_context_v3 +from core.v3 import build_prompt as build_prompt_v3 + + +def build_context( + repo_path: str, + max_files: int = 8, + max_chars: int = 1400, + include_tests: bool = False, +) -> str: + options = BuildOptions( + max_files=max_files, + max_chars=max_chars, + include_tests=include_tests, + ) + context, _ = build_context_v3(repo_path, options) + return context + + +def build_context_with_files( + repo_path: str, + max_files: int = 8, + max_chars: int = 1400, + include_tests: bool = False, +) -> tuple[str, list[str]]: + options = BuildOptions( + max_files=max_files, + max_chars=max_chars, + include_tests=include_tests, + ) + return build_context_v3(repo_path, options) + + +def build_prompt(context: str, question: str, selected_files: list[str] | None = None) -> str: + return build_prompt_v3(context, question, selected_files) diff --git a/core/parser.py b/core/parser.py index c879fd6..da47c11 100644 --- a/core/parser.py +++ b/core/parser.py @@ -1,38 +1,49 @@ import os +from pathlib import Path +from typing import Iterable -IGNORE_DIRS = { +DEFAULT_IGNORE_DIRS = { ".venv", "__pycache__", ".git", "node_modules", - "dist", "build", - "apps", - "frontend", - "infrastructure" + "dist", "build", "apps", "frontend", "infrastructure", "assets" } -EXTENSIONS = (".py", ".js", ".ts", ".tsx") +DEFAULT_EXTENSIONS = (".py", ".js", ".ts", ".tsx", ".rs", ".md") -def get_code_files(repo_path): - code_files = [] +def _iter_code_files( + repo_path: str, + ignore_dirs: set[str], + extensions: tuple[str, ...], +) -> Iterable[Path]: + root_path = Path(repo_path).resolve() - for root, dirs, files in os.walk(repo_path): - dirs[:] = [d for d in dirs if d not in IGNORE_DIRS] + for root, dirs, files in os.walk(root_path): + dirs[:] = sorted(d for d in dirs if d not in ignore_dirs) - for file in files: - if file.endswith(EXTENSIONS): - code_files.append(os.path.join(root, file)) + for filename in sorted(files): + if filename.endswith(extensions): + yield Path(root) / filename - return code_files +def get_code_files( + repo_path: str, + ignore_dirs: set[str] | None = None, + extensions: tuple[str, ...] | None = None, +) -> list[str]: + ignore = ignore_dirs or DEFAULT_IGNORE_DIRS + exts = extensions or DEFAULT_EXTENSIONS + return [str(path) for path in _iter_code_files(repo_path, ignore, exts)] -def read_files(files, max_chars=800): - contents = [] + +def read_files(files: list[str], max_chars: int = 1200) -> str: + contents: list[str] = [] for file in files: try: - with open(file, "r", encoding="utf-8") as f: - text = f.read()[:max_chars] - contents.append(f"\n# FILE: {file}\n{text}") - except: + with open(file, "r", encoding="utf-8") as source: + text = source.read()[:max_chars] + contents.append(f"\n# FILE: {file}\n{text}") + except (OSError, UnicodeDecodeError): continue - return "\n\n".join(contents) \ No newline at end of file + return "\n\n".join(contents) diff --git a/core/v3/__init__.py b/core/v3/__init__.py new file mode 100644 index 0000000..bfcb0f5 --- /dev/null +++ b/core/v3/__init__.py @@ -0,0 +1,4 @@ +from core.v3.context_builder import BuildOptions, build_context, rank_files, score_file +from core.v3.prompting import build_prompt + +__all__ = ["BuildOptions", "build_context", "rank_files", "score_file", "build_prompt"] diff --git a/core/v3/context_builder.py b/core/v3/context_builder.py new file mode 100644 index 0000000..e605284 --- /dev/null +++ b/core/v3/context_builder.py @@ -0,0 +1,66 @@ +from __future__ import annotations + +from dataclasses import dataclass +from pathlib import Path + +from core.parser import get_code_files, read_files + + +KEYWORD_WEIGHTS = { + "main": 6, + "app": 6, + "engine": 7, + "cli": 5, + "api": 5, + "router": 4, + "agent": 4, + "provider": 4, + "pipeline": 3, + "config": 2, + "test": 1, +} + + +@dataclass(frozen=True) +class BuildOptions: + max_files: int = 8 + max_chars: int = 1400 + include_tests: bool = False + + +def score_file(path: str) -> int: + lowered = path.lower() + score = 0 + + for keyword, weight in KEYWORD_WEIGHTS.items(): + if keyword in lowered: + score += weight + + suffix = Path(lowered).suffix + if suffix in {".py", ".rs"}: + score += 2 + elif suffix in {".ts", ".tsx", ".js"}: + score += 1 + + return score + + +def rank_files(files: list[str], include_tests: bool = False) -> list[str]: + ranked: list[tuple[int, str]] = [] + + for file in files: + if not include_tests and "test" in file.lower(): + continue + ranked.append((score_file(file), file)) + + ranked.sort(key=lambda item: (item[0], item[1]), reverse=True) + return [file for _, file in ranked] + + +def build_context(repo_path: str, options: BuildOptions | None = None) -> tuple[str, list[str]]: + opts = options or BuildOptions() + files = get_code_files(repo_path) + ranked_files = rank_files(files, include_tests=opts.include_tests) + selected = ranked_files[: opts.max_files] + context = read_files(selected, max_chars=opts.max_chars) + return context, selected diff --git a/core/v3/prompting.py b/core/v3/prompting.py new file mode 100644 index 0000000..2974306 --- /dev/null +++ b/core/v3/prompting.py @@ -0,0 +1,41 @@ +from __future__ import annotations + + +def build_prompt(context: str, question: str, selected_files: list[str] | None = None) -> str: + selected = "\n".join(f"- {path}" for path in (selected_files or [])) + + return f""" +You are DevOS AI v3, a senior software engineer analyzing a real codebase. + +STRICT RULES: +- Use ONLY the given code context. +- Do NOT invent features or architecture. +- Distinguish facts from inference. +- If context is incomplete, say what is missing. + +SELECTED FILES: +{selected} + +CONTEXT: +{context} + +TASK: +{question} + +Respond in this format: + +1. Purpose +- What this system actually does. + +2. Architecture +- Real modules, boundaries, and data flow. + +3. Key Components +- Explain important files/classes/functions. + +4. Execution Flow +- Step-by-step behavior from input to output. + +5. Risks and Gaps +- Missing validation, edge cases, and technical debt. +""".strip() diff --git a/tests/test_v3_context.py b/tests/test_v3_context.py new file mode 100644 index 0000000..c906391 --- /dev/null +++ b/tests/test_v3_context.py @@ -0,0 +1,29 @@ +import sys +from pathlib import Path + +sys.path.append(str(Path(__file__).resolve().parents[1])) + +from core.v3.context_builder import rank_files, score_file + + +def test_score_file_prioritizes_engine_and_python(): + assert score_file("/tmp/core/engine.py") > score_file("/tmp/docs/readme.md") + + +def test_rank_files_excludes_tests_by_default(): + files = [ + "repo/core/engine.py", + "repo/tests/test_engine.py", + "repo/cli/main.py", + ] + ranked = rank_files(files) + assert "repo/tests/test_engine.py" not in ranked + + +def test_rank_files_can_include_tests(): + files = [ + "repo/core/engine.py", + "repo/tests/test_engine.py", + ] + ranked = rank_files(files, include_tests=True) + assert "repo/tests/test_engine.py" in ranked