Frequently asked questions about DocForge — the document ingest platform that converts PDF, HTML, DOCX, and text into clean Markdown and deterministic RAG chunks for AI agents and vector pipelines.
| Question | Answer |
|---|---|
| What is DocForge? | Document ingest API + MCP server for RAG: file → Markdown + chunked JSON |
| Supported formats? | PDF, HTML, HTM, DOCX, TXT, MD |
| How to self-register (REST)? | POST /api/v1/agents/register with { "name": "…" } |
| How to self-register (MCP)? | Connect docforge-mcp, call register_agent |
| Cursor MCP setup? | Add docforge-mcp to MCP config with DOCFORGE_API_URL |
| Headings vs size chunking? | Headings = section-aware; size = fixed token windows |
| Is chunking deterministic? | Yes — same input + config → same chunk IDs and boundaries |
| API key security? | Store once, never commit; header X-DocForge-Key |
| RAG pipeline usage? | Ingest → embed chunks by id → upsert to vector store |
| Server not running? | Run ./scripts/dev.sh or docforge-api on port 8787 |
| Unsupported format? | Use PDF, HTML, DOCX, TXT, or MD only |
| Duplicate agent name? | Pick a unique name or reuse your existing key |
| Production use? | Stateless ingest, API keys, Docker; scale behind your infra |
Canonical URLs: /docs/faq.md · /FAQ.md · /api/v1/agent-manifest · /AGENTS.md · /llms.txt
DocForge is an open document ingest platform for RAG (Retrieval-Augmented Generation) pipelines and AI agents. It accepts PDF, HTML, DOCX, and plain-text files, converts them to normalized Markdown, and produces deterministic chunked JSON with token estimates. Agents discover DocForge via GET /api/v1/agent-manifest or /AGENTS.md, self-register for an API key, and ingest documents through REST (POST /api/v1/ingest) or MCP (ingest_document tool). DocForge is designed for vector-store ingestion, agent document processing, and enterprise RAG workflows where reproducible chunk boundaries matter.
DocForge supports these input formats:
| Extension | Format | Conversion |
|---|---|---|
.pdf |
Text extraction via PyMuPDF | |
.html, .htm |
HTML | HTML → clean Markdown |
.docx |
Microsoft Word | DOCX → Markdown |
.txt |
Plain text | Passed through with normalization |
.md, .markdown |
Markdown | Normalized whitespace and structure |
Unsupported formats (e.g. .xlsx, .pptx, images-only PDFs) return an error. Use /api/v1/convert to get Markdown only, or /api/v1/ingest for the full Markdown + chunks pipeline.
DocForge splits converted Markdown into chunks using one of two strategies: headings (split on H1–H6 with optional sub-splitting by token size) or size (fixed token windows with overlap). Each chunk includes id, index, content, token_estimate, char_count, and metadata. Chunk IDs are derived from SHA256 of document_id:index:content, so re-ingesting the same document with the same config produces identical chunk IDs — ideal for idempotent vector upserts.
Default chunk config: strategy=headings, max_tokens=512, overlap_tokens=64, token_model=cl100k_base.
Token estimates are computed with tiktoken using the cl100k_base encoding by default (compatible with OpenAI GPT-4 and text-embedding-ada-002 family models). DocForge reports token_estimate on the full document and on each chunk. Estimates are consistent within DocForge but may differ slightly from a cloud provider's billing tokenizer. Configure token_model in chunk requests to change the encoding.
| Aspect | headings |
size |
|---|---|---|
| Split logic | Markdown heading levels (H1–H6) | Fixed token count windows |
| Metadata | heading_path array per chunk |
Overlap between consecutive chunks |
| Best for | Manuals, wikis, structured reports | Logs, transcripts, unstructured text |
| Oversized sections | Sub-split by max_tokens |
N/A — windows are token-based |
| Overlap | None between sections | Configurable overlap_tokens (default 64) |
Both strategies are fully deterministic: identical Markdown + config always yields the same chunks.
Use headings when your source documents have clear section structure — PDF reports, HTML pages, DOCX policies, technical manuals. Chunks align with semantic sections and include a heading_path like ["Introduction", "Installation"] for better retrieval context.
Use size when structure is weak or absent — chat logs, transcripts, plain-text dumps, or when you need fixed embedding window sizes (e.g. 256–512 tokens). Set overlap_tokens to preserve context across chunk boundaries.
See /docs/agents/chunking.md for parameter details and RAG recommendations by source type.
- Discover —
GET /api/v1/agent-manifestor read/AGENTS.md. - Register — send a POST request:
POST /api/v1/agents/register
Content-Type: application/json
{
"name": "my-rag-ingest-bot",
"description": "Production document pipeline",
"capabilities": ["ingest", "chunk", "convert"]
}- Save the response — the JSON includes
id,api_key(shown once),cursor_config, andmcp_env. - Authenticate — send
X-DocForge-Key: df_…on subsequent REST calls (optional but recommended for tracking).
Example with curl:
curl -sS -X POST 'http://127.0.0.1:8787/api/v1/agents/register' \
-H 'content-type: application/json' \
-d '{"name":"my-rag-agent","capabilities":["ingest","chunk"]}'- Install DocForge (
pip install -e .) sodocforge-mcpis on PATH. - Connect MCP with command
docforge-mcpand envDOCFORGE_API_URL=http://127.0.0.1:8787(or your deployment URL). - Optionally call
get_agent_manifestto discover endpoints and tools. - Call
register_agentwith{ "name": "your-agent-name", "description": "optional" }. - Save
api_keyfrom the JSON result — it cannot be retrieved again. - Set
DOCFORGE_API_KEYin your MCP config env for authenticated calls (agent_whoami, tracked ingest).
After registration, use ingest_document, convert_document, and chunk_markdown tools.
- Self-register via REST or MCP and save your
api_key. - Add to Cursor MCP settings (merge
cursor_configfrom the register response):
{
"mcpServers": {
"docforge": {
"command": "docforge-mcp",
"env": {
"DOCFORGE_API_URL": "http://127.0.0.1:8787",
"DOCFORGE_API_KEY": "df_YOUR_KEY_FROM_REGISTER"
}
}
}
}- Restart Cursor or reload MCP servers.
- Verify with the
health_checkoragent_whoamitool.
For local development, ensure DocForge API is running (./scripts/dev.sh). For production, set DOCFORGE_API_URL to your deployed origin (e.g. https://docforge.example.com).
See /docs/agents/mcp.md for the full tool reference.
DocForge API keys use the format df_<urlsafe-secret>. Obtain a key by self-registering an agent. Send it on REST requests via the header:
X-DocForge-Key: df_YOUR_KEYFor MCP, set the environment variable DOCFORGE_API_KEY. The key is optional for ingest/convert/chunk but recommended for usage tracking and agent_whoami. Keys are shown only once at registration and cannot be recovered — store them in a secret manager or MCP env, not in git.
- Register once per agent identity — do not re-register repeatedly.
- Never commit, log, or paste
api_keyinto issues, chat transcripts, or public repos. - Store keys in MCP env,
.env(gitignored), or a host secret store. - Use
agent_whoamito verify a key without exposing it in output. - Treat chunk output as potentially sensitive document content per your data classification policy.
- Rotate by registering a new agent identity if a key is compromised (DocForge does not expose key rotation APIs).
Yes. DocForge chunking is fully deterministic: the same file content, conversion output, and chunk configuration always produce identical chunk boundaries, content, and id values. Chunk IDs are the first 16 hex characters of SHA256(document_id:index:content). Use these IDs for idempotent vector store upserts — re-ingesting unchanged documents will not create duplicate vectors when keyed by chunk id.
A typical RAG pipeline with DocForge:
- Ingest —
POST /api/v1/ingestor MCPingest_documentwith your file and chunk strategy. - Receive — response includes
markdown,chunks[],document_id, and token totals. - Embed — send each chunk's
contentto your embedding model (OpenAI, Cohere, local, etc.). - Store — upsert vectors keyed by chunk
idinto Pinecone, Weaviate, pgvector, Chroma, or similar. - Retrieve — at query time, search vectors and pass matched
content+heading_pathto your LLM.
Prefer ingest_document (full pipeline) over separate convert + chunk calls. Choose headings for structured docs and size for unstructured text. See /docs/agents/chunking.md for strategy recommendations.
The agent manifest is machine-readable JSON at GET /api/v1/agent-manifest (also GET /.well-known/docforge.json). It includes:
origin,version, and descriptiondocumentationlinks (AGENTS.md, llms.txt, MCP/REST/chunking guides, FAQ)selfRegistrationREST and MCP instructionsmcp.toolslist with input schemasrest.endpointswith auth requirementssupportedFormats,chunkingdefaults, and curl examplescursorConfigtemplate for MCP setup
Agents should fetch the manifest first to discover all URLs and capabilities without hardcoding paths.
| Tool | Auth | Purpose |
|---|---|---|
get_agent_manifest |
None | Fetch manifest JSON |
register_agent |
None | Self-register; returns api_key once |
list_agents |
None | List agents (no secrets) |
health_check |
None | API health + version |
convert_document |
Optional key | Local file → Markdown |
chunk_markdown |
Optional key | Markdown → deterministic chunks |
ingest_document |
Optional key | Full pipeline: file → MD + chunks |
agent_whoami |
Required key | Your registered profile |
Transport: stdio via command docforge-mcp. See /docs/agents/mcp.md.
| Method | Path | Auth | Purpose |
|---|---|---|---|
| GET | /api/v1/agent-manifest |
— | Agent discovery JSON |
| GET | /.well-known/docforge.json |
— | Same manifest |
| GET | /AGENTS.md |
— | Agent guide (Markdown) |
| GET | /llms.txt |
— | LLM discovery index |
| GET | /docs/faq.md |
— | This FAQ |
| POST | /api/v1/agents/register |
— | Self-register agent |
| GET | /api/v1/agents |
— | List agents |
| GET | /api/v1/agents/me |
Key | Authenticated profile |
| POST | /api/v1/ingest |
Optional | File → Markdown + chunks |
| POST | /api/v1/convert |
Optional | File → Markdown |
| POST | /api/v1/chunk |
Optional | Markdown → chunks |
| GET | /api/v1/health |
— | Health check |
Interactive OpenAPI docs: /api/docs. Full reference: /docs/agents/rest.md.
| Context | DOCFORGE_API_URL / origin |
Notes |
|---|---|---|
| Local dev | http://127.0.0.1:8787 |
Default from ./scripts/dev.sh |
| Docker | http://localhost:8787 or service name |
See docker compose config |
| Production | https://your-domain.com |
Set in MCP env and agent manifest |
The register response includes cursor_config and mcp_env with the correct origin for your deployment. Always use the manifest's origin field rather than hardcoding localhost in production. CORS is enabled for browser and agent access.
If REST calls fail with connection refused or MCP health_check reports offline:
- Start the API:
./scripts/dev.sh(recommended) ordocforge-apiafterpip install -e ".[dev]". - Verify:
curl http://127.0.0.1:8787/api/v1/health— expect{"status":"ok",…}. - Check port — default is 8787 (
DOCFORGE_PORTin.env). - For MCP, ensure
DOCFORGE_API_URLmatches the running server (not a stale URL). - Rebuild the web UI if needed:
cd web && npm run build.
The human UI at / also shows version status in the header.
DocForge only accepts PDF, HTML, HTM, DOCX, TXT, and Markdown. If you see an unsupported format error:
- Confirm the file extension matches a supported type.
- Convert exotic formats externally (e.g. PPTX → PDF, CSV → TXT) before ingest.
- For scanned PDFs with no text layer, run OCR first — DocForge extracts embedded text only via PyMuPDF.
- Check
supportedFormatsin/api/v1/agent-manifestfor the authoritative list.
Use the UI upload dialog or REST multipart field file=@document.pdf.
Registration returns HTTP 409 with "Agent name already registered" when the agent name is taken. DocForge agent names must be unique.
Fix options:
- Reuse your existing key — if you already registered that name, use the saved
api_key; do not re-register. - Choose a unique name — e.g.
my-rag-bot-prod-2026instead ofmy-rag-bot. - List agents —
GET /api/v1/agentsor MCPlist_agentsto see registered names (IDs only, no keys).
API keys cannot be retrieved after registration. If you lost a key, register a new agent with a different name.
DocForge is designed for production RAG ingest with these characteristics:
- Stateless document processing — convert and chunk in request/response; no mandatory cloud dependency.
- Agent authentication — API keys via
X-DocForge-Keyfor tracking and access control patterns. - Deterministic output — reproducible chunks for auditable pipelines.
- Docker deployment —
docker compose up --buildfor containerized runs. - Configurable limits —
DOCFORGE_MAX_UPLOAD_MB(default 50 MB), CORS, host/port via environment. - OpenAPI —
/api/docsfor integration testing and client generation.
For enterprise deployments, run DocForge behind your API gateway, terminate TLS at the load balancer, store keys in a secret manager, and apply your org's data retention policy to uploaded documents and chunk output.
REST:
curl -X POST http://127.0.0.1:8787/api/v1/ingest \
-H 'X-DocForge-Key: df_YOUR_KEY' \
-F 'file=@report.pdf' \
-F 'strategy=headings' \
-F 'max_tokens=512'MCP: call ingest_document with { "file_path": "/path/to/report.pdf", "strategy": "headings", "max_tokens": 512 }.
UI: open / → upload PDF → configure chunking → Ingest → Markdown + Chunks.
Response includes markdown, chunks, document_id, chunk_count, and total_chunk_tokens.
Yes. Use POST /api/v1/convert (REST) or MCP convert_document. These return clean Markdown and token estimates without producing chunks. To chunk existing Markdown separately, use POST /api/v1/chunk or MCP chunk_markdown. For most RAG workflows, ingest (convert + chunk in one call) is simpler and ensures consistent document_id across chunks.
Each chunk's id is the first 16 hexadecimal characters of:
SHA256(document_id + ":" + index + ":" + content)
The document_id is assigned per ingest request. Because the hash includes chunk content and index, IDs are stable across re-runs with identical input and config. Use chunk id as the primary key for vector store upserts to avoid duplicates on re-ingestion.
| Resource | URL |
|---|---|
| Agent guide | /AGENTS.md |
| LLM index | /llms.txt |
| MCP reference | /docs/agents/mcp.md |
| REST reference | /docs/agents/rest.md |
| Chunking guide | /docs/agents/chunking.md |
| Cursor skill | /docs/skill.md |
| Agent manifest | /api/v1/agent-manifest |
| OpenAPI | /api/docs |