Backend API for video or audio transcription using whisper.
Note:
This project is mostly a proof-of-concept around whisper speech to text capabilities
- Runtime: Node.js 24 + TypeScript
- Framework: Express + Socket.io
- Database: PostgreSQL 17 + Prisma ORM
- Transcription: whisper-cpp-server
- Summarization: OpenRouter API
- Docker & Docker Compose
- An OpenRouter API key
Clone the project from GitLab:
git clone https://gitlab.isaid.fr/Balanced436/insights.gitCreate a secrets/ directory at the root of the project containing the following files:
| File | Description |
|---|---|
secrets/postgres_db.txt |
PostgreSQL database name |
secrets/postgres_password.txt |
PostgreSQL password |
secrets/postgres_user.txt |
PostgreSQL username |
secrets/open_router_api_key.txt |
OpenRouter API key |
Download the Whisper model from Hugging Face:
cd whisper
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
python3 download-model.pyStart the containers with:
docker compose up -dOnce running, the services (whisper-cpp, insights-db, and insights-backend) are accessible, with the API server listening on port 4000.
Before running a transcription, create a corpus to group related audio/video sources. Send a POST request to /corpus:
curl --request POST 'http://localhost:4000/corpus' \
--header 'Content-Type: application/json' \
--data '{
"title" : "Public Sessions",
"description" : "Corpus containing public assembly sessions."
}'The server responds with the created corpus object and its ID:
{
"corpus": {
"id": 1,
"description": "Corpus containing public assembly sessions.",
"title": "Public Sessions",
"createdAt": "2026-08-21T19:42:02.916Z",
"updatedAt": "2026-08-21T19:42:02.916Z"
}
}Upload an audio file to a corpus by sending a POST request to /source:
curl 'http://localhost:4000/source' \
--form 'audio=@"session_audio.wav";type=audio/wav' \
--form 'title=Demo Source' \
--form 'description=Public session recording' \
--form 'corpusID=1'The server confirms the creation of the source:
{
"message": "source created successfully",
"data": {
"id": 1,
"title": "Demo Source",
"description": "Public session recording",
"videoUrl": null,
"audioUrl": "/app/source/audio/1788078343584.wav",
"createdAt": "2026-08-22T10:09:03.476Z",
"updatedAt": "2026-08-22T10:09:03.476Z",
"corpusID": 1
}
}Start the Whisper transcription process by sending a POST request to /transcription:
curl --request POST 'http://localhost:4000/transcription' \
--header 'Content-Type: application/json' \
--data '{"sourceId": 1}'The server returns a task object with status PENDING:
{
"task": {
"id": 1,
"type": "TRANSCRIPTION",
"createdAt": "2026-08-22T14:30:22.591Z",
"finishedAt": null,
"status": "PENDING",
"transcriptionId": 1
}
}Track the status of an ongoing task by polling the /task endpoint:
curl 'http://localhost:4000/task'{
"data": [
{
"id": 1,
"type": "TRANSCRIPTION",
"createdAt": "2026-08-22T14:30:22.591Z",
"finishedAt": "2026-08-22T14:31:28.575Z",
"status": "COMPLETED",
"transcriptionId": 1
}
]
}Once the status is COMPLETED, retrieve the transcribed text with a GET request to /transcription:
curl 'http://localhost:4000/transcription'[
{
"id": 1,
"content": "Transcribed audio content...",
"createdAt": "2026-08-22T14:30:22.579Z",
"updatedAt": "2026-08-22T14:31:28.570Z",
"sourceId": 1
}
]Generate an AI summary of a transcription using OpenRouter by sending a POST request to /summary:
curl --request POST 'http://localhost:4000/summary' \
--header 'Content-Type: application/json' \
--data '{"transcriptionId": 1}'The server returns a task object of type SUMMARIZATION with status PENDING:
{
"task": {
"id": 2,
"type": "SUMMARIZATION",
"createdAt": "2026-08-23T06:42:07.106Z",
"finishedAt": null,
"status": "PENDING",
"transcriptionId": 1
}
}To fetch the generated summary, send a GET request to /summary/:id:
curl 'http://localhost:4000/summary/1'{
"data": [
{
"id": 1,
"content": "Summary of the transcription content...",
"createdAt": "2026-08-23T06:42:07.152Z",
"updatedAt": "2026-08-23T06:42:14.052Z",
"transcriptionId": 1
}
]
}The Corpus resource represents a thematic collection of sources.
Retrieve one or all corpora. Omitting id returns all.
curl http://localhost:4000/corpus # all
curl http://localhost:4000/corpus/1 # singleCreate a corpus. title is required, description optional.
curl -X POST http://localhost:4000/corpus \
-H "Content-Type: application/json" \
-d '{"title": "Conférences 2025", "description": "Collection de conférences techniques"}'Update an existing corpus (only provided fields).
curl -X PATCH http://localhost:4000/corpus/1 \
-H "Content-Type: application/json" \
-d '{"title": "Conférences 2025 (Mis à jour)"}'Delete a corpus.
curl -X DELETE http://localhost:4000/corpus/1A Source is a media file (video/audio) belonging to a Corpus.
Retrieve one or all sources. Supports ?corpusid= filter.
curl http://localhost:4000/source # all
curl http://localhost:4000/source/1 # single
curl http://localhost:4000/source?corpusid=1 # by corpusCreate a source with a file upload. Accepts multipart/form-data.
curl -X POST http://localhost:4000/source \
-F "video=@/path/to/video.mp4" \
-F "title=Keynote" \
-F "description=Opening keynote" \
-F "corpusID=1"Update a source
curl -X PUT http://localhost:4000/source/1 \
-F "title=Keynote (Updated)" \
-F "video=@/path/to/new-video.mp4"Delete a source and its associated files.
curl -X DELETE http://localhost:4000/source/1A Transcription is the text result of transcribing a Source's audio file.
Retrieve one or all transcriptions. Supports ?sourceid= filter.
curl http://localhost:4000/transcription # all
curl http://localhost:4000/transcription/1 # single
curl http://localhost:4000/transcription?sourceid=1 # by sourceCreate a transcription for a source. sourceId is required. Optionally set skipTranscription to true to skip actual Whisper inference and use placeholder text.
curl -X POST http://localhost:4000/transcription \
-H "Content-Type: application/json" \
-d '{"sourceId": 1}'Update a transcription's content.
curl -X PUT http://localhost:4000/transcription/1 \
-H "Content-Type: application/json" \
-d '{"content": "Revised transcription text"}'Delete a transcription.
curl -X DELETE http://localhost:4000/transcription/1A Summary is the generated summary of a Transcription's content, produced via the OpenRouter API.
Retrieve one or all summaries. Supports ?transcriptionid= and ?sourceid= filters.
curl http://localhost:4000/summary # all
curl http://localhost:4000/summary/1 # single
curl http://localhost:4000/summary?transcriptionid=1 # by transcription
curl http://localhost:4000/summary?sourceid=1 # by sourceCreate a summary for a transcription. transcriptionId is required. Optionally pass content to set the summary directly instead of generating it via OpenRouter.
curl -X POST http://localhost:4000/summary \
-H "Content-Type: application/json" \
-d '{"transcriptionId": 1}'Delete a summary.
curl -X DELETE http://localhost:4000/summary/1A Task represents an asynchronous job (transcription or summarization) with a PENDING, COMPLETED, or ERROR status. Tasks are emitted in real-time via Socket.io.
Retrieve one or all tasks.
curl http://localhost:4000/task # all
curl http://localhost:4000/task/1 # singleCreate a task. transcriptionId and taskType are required.
curl -X POST http://localhost:4000/task \
-H "Content-Type: application/json" \
-d '{"transcriptionId": 1, "taskType": "TRANSCRIPTION"}'Update a task's status.
curl -X PUT http://localhost:4000/task/1 \
-H "Content-Type: application/json" \
-d '{"status": "COMPLETED"}'