Skip to content

Latest commit

 

History

History
371 lines (252 loc) · 17.4 KB

File metadata and controls

371 lines (252 loc) · 17.4 KB

DocForge FAQ

Frequently asked questions about DocForge — the document ingest platform that converts PDF, HTML, DOCX, and text into clean Markdown and deterministic RAG chunks for AI agents and vector pipelines.

Quick answers

Question Answer
What is DocForge? Document ingest API + MCP server for RAG: file → Markdown + chunked JSON
Supported formats? PDF, HTML, HTM, DOCX, TXT, MD
How to self-register (REST)? POST /api/v1/agents/register with { "name": "…" }
How to self-register (MCP)? Connect docforge-mcp, call register_agent
Cursor MCP setup? Add docforge-mcp to MCP config with DOCFORGE_API_URL
Headings vs size chunking? Headings = section-aware; size = fixed token windows
Is chunking deterministic? Yes — same input + config → same chunk IDs and boundaries
API key security? Store once, never commit; header X-DocForge-Key
RAG pipeline usage? Ingest → embed chunks by id → upsert to vector store
Server not running? Run ./scripts/dev.sh or docforge-api on port 8787
Unsupported format? Use PDF, HTML, DOCX, TXT, or MD only
Duplicate agent name? Pick a unique name or reuse your existing key
Production use? Stateless ingest, API keys, Docker; scale behind your infra

Canonical URLs: /docs/faq.md · /FAQ.md · /api/v1/agent-manifest · /AGENTS.md · /llms.txt


What is DocForge?

DocForge is an open document ingest platform for RAG (Retrieval-Augmented Generation) pipelines and AI agents. It accepts PDF, HTML, DOCX, and plain-text files, converts them to normalized Markdown, and produces deterministic chunked JSON with token estimates. Agents discover DocForge via GET /api/v1/agent-manifest or /AGENTS.md, self-register for an API key, and ingest documents through REST (POST /api/v1/ingest) or MCP (ingest_document tool). DocForge is designed for vector-store ingestion, agent document processing, and enterprise RAG workflows where reproducible chunk boundaries matter.


What document formats does DocForge support?

DocForge supports these input formats:

Extension Format Conversion
.pdf PDF Text extraction via PyMuPDF
.html, .htm HTML HTML → clean Markdown
.docx Microsoft Word DOCX → Markdown
.txt Plain text Passed through with normalization
.md, .markdown Markdown Normalized whitespace and structure

Unsupported formats (e.g. .xlsx, .pptx, images-only PDFs) return an error. Use /api/v1/convert to get Markdown only, or /api/v1/ingest for the full Markdown + chunks pipeline.


How does DocForge chunking work?

DocForge splits converted Markdown into chunks using one of two strategies: headings (split on H1–H6 with optional sub-splitting by token size) or size (fixed token windows with overlap). Each chunk includes id, index, content, token_estimate, char_count, and metadata. Chunk IDs are derived from SHA256 of document_id:index:content, so re-ingesting the same document with the same config produces identical chunk IDs — ideal for idempotent vector upserts.

Default chunk config: strategy=headings, max_tokens=512, overlap_tokens=64, token_model=cl100k_base.


What are token estimates in DocForge?

Token estimates are computed with tiktoken using the cl100k_base encoding by default (compatible with OpenAI GPT-4 and text-embedding-ada-002 family models). DocForge reports token_estimate on the full document and on each chunk. Estimates are consistent within DocForge but may differ slightly from a cloud provider's billing tokenizer. Configure token_model in chunk requests to change the encoding.


What is the difference between headings and size chunking?

Aspect headings size
Split logic Markdown heading levels (H1–H6) Fixed token count windows
Metadata heading_path array per chunk Overlap between consecutive chunks
Best for Manuals, wikis, structured reports Logs, transcripts, unstructured text
Oversized sections Sub-split by max_tokens N/A — windows are token-based
Overlap None between sections Configurable overlap_tokens (default 64)

Both strategies are fully deterministic: identical Markdown + config always yields the same chunks.


When should I use headings vs size chunking?

Use headings when your source documents have clear section structure — PDF reports, HTML pages, DOCX policies, technical manuals. Chunks align with semantic sections and include a heading_path like ["Introduction", "Installation"] for better retrieval context.

Use size when structure is weak or absent — chat logs, transcripts, plain-text dumps, or when you need fixed embedding window sizes (e.g. 256–512 tokens). Set overlap_tokens to preserve context across chunk boundaries.

See /docs/agents/chunking.md for parameter details and RAG recommendations by source type.


How do AI agents self-register with DocForge via REST?

  1. Discover — GET /api/v1/agent-manifest or read /AGENTS.md.
  2. Register — send a POST request:
POST /api/v1/agents/register
Content-Type: application/json

{
  "name": "my-rag-ingest-bot",
  "description": "Production document pipeline",
  "capabilities": ["ingest", "chunk", "convert"]
}
  1. Save the response — the JSON includes id, api_key (shown once), cursor_config, and mcp_env.
  2. Authenticate — send X-DocForge-Key: df_… on subsequent REST calls (optional but recommended for tracking).

Example with curl:

curl -sS -X POST 'http://127.0.0.1:8787/api/v1/agents/register' \
  -H 'content-type: application/json' \
  -d '{"name":"my-rag-agent","capabilities":["ingest","chunk"]}'

How do AI agents self-register with DocForge via MCP?

  1. Install DocForge (pip install -e .) so docforge-mcp is on PATH.
  2. Connect MCP with command docforge-mcp and env DOCFORGE_API_URL=http://127.0.0.1:8787 (or your deployment URL).
  3. Optionally call get_agent_manifest to discover endpoints and tools.
  4. Call register_agent with { "name": "your-agent-name", "description": "optional" }.
  5. Save api_key from the JSON result — it cannot be retrieved again.
  6. Set DOCFORGE_API_KEY in your MCP config env for authenticated calls (agent_whoami, tracked ingest).

After registration, use ingest_document, convert_document, and chunk_markdown tools.


How do I set up DocForge MCP in Cursor?

  1. Self-register via REST or MCP and save your api_key.
  2. Add to Cursor MCP settings (merge cursor_config from the register response):
{
  "mcpServers": {
    "docforge": {
      "command": "docforge-mcp",
      "env": {
        "DOCFORGE_API_URL": "http://127.0.0.1:8787",
        "DOCFORGE_API_KEY": "df_YOUR_KEY_FROM_REGISTER"
      }
    }
  }
}
  1. Restart Cursor or reload MCP servers.
  2. Verify with the health_check or agent_whoami tool.

For local development, ensure DocForge API is running (./scripts/dev.sh). For production, set DOCFORGE_API_URL to your deployed origin (e.g. https://docforge.example.com).

See /docs/agents/mcp.md for the full tool reference.


What is the DocForge API key format and how do I use it?

DocForge API keys use the format df_<urlsafe-secret>. Obtain a key by self-registering an agent. Send it on REST requests via the header:

X-DocForge-Key: df_YOUR_KEY

For MCP, set the environment variable DOCFORGE_API_KEY. The key is optional for ingest/convert/chunk but recommended for usage tracking and agent_whoami. Keys are shown only once at registration and cannot be recovered — store them in a secret manager or MCP env, not in git.


How do I keep DocForge API keys secure?

  • Register once per agent identity — do not re-register repeatedly.
  • Never commit, log, or paste api_key into issues, chat transcripts, or public repos.
  • Store keys in MCP env, .env (gitignored), or a host secret store.
  • Use agent_whoami to verify a key without exposing it in output.
  • Treat chunk output as potentially sensitive document content per your data classification policy.
  • Rotate by registering a new agent identity if a key is compromised (DocForge does not expose key rotation APIs).

Is DocForge chunking deterministic?

Yes. DocForge chunking is fully deterministic: the same file content, conversion output, and chunk configuration always produce identical chunk boundaries, content, and id values. Chunk IDs are the first 16 hex characters of SHA256(document_id:index:content). Use these IDs for idempotent vector store upserts — re-ingesting unchanged documents will not create duplicate vectors when keyed by chunk id.


How do I use DocForge in a RAG pipeline?

A typical RAG pipeline with DocForge:

  1. Ingest — POST /api/v1/ingest or MCP ingest_document with your file and chunk strategy.
  2. Receive — response includes markdown, chunks[], document_id, and token totals.
  3. Embed — send each chunk's content to your embedding model (OpenAI, Cohere, local, etc.).
  4. Store — upsert vectors keyed by chunk id into Pinecone, Weaviate, pgvector, Chroma, or similar.
  5. Retrieve — at query time, search vectors and pass matched content + heading_path to your LLM.

Prefer ingest_document (full pipeline) over separate convert + chunk calls. Choose headings for structured docs and size for unstructured text. See /docs/agents/chunking.md for strategy recommendations.


What is the DocForge agent manifest?

The agent manifest is machine-readable JSON at GET /api/v1/agent-manifest (also GET /.well-known/docforge.json). It includes:

  • origin, version, and description
  • documentation links (AGENTS.md, llms.txt, MCP/REST/chunking guides, FAQ)
  • selfRegistration REST and MCP instructions
  • mcp.tools list with input schemas
  • rest.endpoints with auth requirements
  • supportedFormats, chunking defaults, and curl examples
  • cursorConfig template for MCP setup

Agents should fetch the manifest first to discover all URLs and capabilities without hardcoding paths.


What MCP tools does DocForge provide?

Tool Auth Purpose
get_agent_manifest None Fetch manifest JSON
register_agent None Self-register; returns api_key once
list_agents None List agents (no secrets)
health_check None API health + version
convert_document Optional key Local file → Markdown
chunk_markdown Optional key Markdown → deterministic chunks
ingest_document Optional key Full pipeline: file → MD + chunks
agent_whoami Required key Your registered profile

Transport: stdio via command docforge-mcp. See /docs/agents/mcp.md.


What REST API endpoints does DocForge expose?

Method Path Auth Purpose
GET /api/v1/agent-manifest — Agent discovery JSON
GET /.well-known/docforge.json — Same manifest
GET /AGENTS.md — Agent guide (Markdown)
GET /llms.txt — LLM discovery index
GET /docs/faq.md — This FAQ
POST /api/v1/agents/register — Self-register agent
GET /api/v1/agents — List agents
GET /api/v1/agents/me Key Authenticated profile
POST /api/v1/ingest Optional File → Markdown + chunks
POST /api/v1/convert Optional File → Markdown
POST /api/v1/chunk Optional Markdown → chunks
GET /api/v1/health — Health check

Interactive OpenAPI docs: /api/docs. Full reference: /docs/agents/rest.md.


What is the difference between local and deployed DocForge URLs?

Context DOCFORGE_API_URL / origin Notes
Local dev http://127.0.0.1:8787 Default from ./scripts/dev.sh
Docker http://localhost:8787 or service name See docker compose config
Production https://your-domain.com Set in MCP env and agent manifest

The register response includes cursor_config and mcp_env with the correct origin for your deployment. Always use the manifest's origin field rather than hardcoding localhost in production. CORS is enabled for browser and agent access.


DocForge server not running — how do I fix it?

If REST calls fail with connection refused or MCP health_check reports offline:

  1. Start the API: ./scripts/dev.sh (recommended) or docforge-api after pip install -e ".[dev]".
  2. Verify: curl http://127.0.0.1:8787/api/v1/health — expect {"status":"ok",…}.
  3. Check port — default is 8787 (DOCFORGE_PORT in .env).
  4. For MCP, ensure DOCFORGE_API_URL matches the running server (not a stale URL).
  5. Rebuild the web UI if needed: cd web && npm run build.

The human UI at / also shows version status in the header.


Unsupported file format error — what do I do?

DocForge only accepts PDF, HTML, HTM, DOCX, TXT, and Markdown. If you see an unsupported format error:

  1. Confirm the file extension matches a supported type.
  2. Convert exotic formats externally (e.g. PPTX → PDF, CSV → TXT) before ingest.
  3. For scanned PDFs with no text layer, run OCR first — DocForge extracts embedded text only via PyMuPDF.
  4. Check supportedFormats in /api/v1/agent-manifest for the authoritative list.

Use the UI upload dialog or REST multipart field file=@document.pdf.


Duplicate agent name error — what do I do?

Registration returns HTTP 409 with "Agent name already registered" when the agent name is taken. DocForge agent names must be unique.

Fix options:

  1. Reuse your existing key — if you already registered that name, use the saved api_key; do not re-register.
  2. Choose a unique name — e.g. my-rag-bot-prod-2026 instead of my-rag-bot.
  3. List agents — GET /api/v1/agents or MCP list_agents to see registered names (IDs only, no keys).

API keys cannot be retrieved after registration. If you lost a key, register a new agent with a different name.


Is DocForge suitable for production and enterprise?

DocForge is designed for production RAG ingest with these characteristics:

  • Stateless document processing — convert and chunk in request/response; no mandatory cloud dependency.
  • Agent authentication — API keys via X-DocForge-Key for tracking and access control patterns.
  • Deterministic output — reproducible chunks for auditable pipelines.
  • Docker deployment — docker compose up --build for containerized runs.
  • Configurable limits — DOCFORGE_MAX_UPLOAD_MB (default 50 MB), CORS, host/port via environment.
  • OpenAPI — /api/docs for integration testing and client generation.

For enterprise deployments, run DocForge behind your API gateway, terminate TLS at the load balancer, store keys in a secret manager, and apply your org's data retention policy to uploaded documents and chunk output.


How do I ingest a PDF with DocForge?

REST:

curl -X POST http://127.0.0.1:8787/api/v1/ingest \
  -H 'X-DocForge-Key: df_YOUR_KEY' \
  -F 'file=@report.pdf' \
  -F 'strategy=headings' \
  -F 'max_tokens=512'

MCP: call ingest_document with { "file_path": "/path/to/report.pdf", "strategy": "headings", "max_tokens": 512 }.

UI: open / → upload PDF → configure chunking → Ingest → Markdown + Chunks.

Response includes markdown, chunks, document_id, chunk_count, and total_chunk_tokens.


Can I convert a document to Markdown without chunking?

Yes. Use POST /api/v1/convert (REST) or MCP convert_document. These return clean Markdown and token estimates without producing chunks. To chunk existing Markdown separately, use POST /api/v1/chunk or MCP chunk_markdown. For most RAG workflows, ingest (convert + chunk in one call) is simpler and ensures consistent document_id across chunks.


How are chunk IDs generated?

Each chunk's id is the first 16 hexadecimal characters of:

SHA256(document_id + ":" + index + ":" + content)

The document_id is assigned per ingest request. Because the hash includes chunk content and index, IDs are stable across re-runs with identical input and config. Use chunk id as the primary key for vector store upserts to avoid duplicates on re-ingestion.


Related documentation

Resource URL
Agent guide /AGENTS.md
LLM index /llms.txt
MCP reference /docs/agents/mcp.md
REST reference /docs/agents/rest.md
Chunking guide /docs/agents/chunking.md
Cursor skill /docs/skill.md
Agent manifest /api/v1/agent-manifest
OpenAPI /api/docs