Skip to content

Langfuse traces export inline attachments in full, once per span that serializes the conversation #578

Description

@usnavy13

Summary

When a message carries an inline attachment (image, PDF, audio or video), every Langfuse span whose input or output serializes the conversation exports the full base64 payload. One agent turn exports the same file from the agent root span, the agent span and the llm generation. Runs that go through AgentModelCall also export it from that chain and its prompt span. Each later turn that replays the attachment exports it again.

Observed on @librechat/agents 3.9.7 with @langfuse/otel 5.10.1.

Reproduction

  1. Enable Langfuse tracing with LANGFUSE_MEDIA_UPLOAD_ENABLED=false.
  2. Send a Google turn with a 2.86 MiB PDF attached as { type: 'media', mimeType: 'application/pdf', data: '<base64>' }.
  3. Read the trace's observations.

The trace carries 15.3 MiB. The 3.8 MiB payload appears four times:

Observation Input Output
agent root (<agent name>) 3.8 MiB 3.8 MiB
agent 3.8 MiB —
llm 3.8 MiB —

The same file sent to an OpenAI Responses model as an input_file data: URI produces the same 15.3 MiB trace.

Why the Langfuse SDK does not catch it

LangfuseSpanProcessor extracts only data:<mime>;base64, URIs, and only while media upload is enabled. Everything else is exported as text:

  • Raw base64: Google media and inlineData parts, Anthropic source: { type: 'base64' }, and OpenAI-compatible input_audio.
  • Bedrock bytes buffers, which serialize as {"type":"Buffer","data":[37,80,68,…]}, about 3.5× the file size.
  • data: URIs too, whenever media upload is disabled. Self-hosted instances without direct blob-store access have to disable it.

LangfuseConfig offers no way to rewrite span content. The SDK's mask option is not forwarded. A function-valued option would not fit the processor cache key either, because that key is built with JSON.stringify.

Impact

  • Size scales as file size × spans that serialize the conversation × turns that replay the file. A 20 MB PDF is about 27 MiB of base64 per copy, so a single turn exports more than 100 MiB.
  • An OTLP batch that exceeds a collector's or Langfuse's request-size limit is rejected whole. That drops the small spans batched with it, including timings, usage and cost. Batches that are accepted take memory to ingest and storage to keep.
  • In the host process, each span serializes its own copy when it starts (about 35 ms per 27 MiB on Node 24) and holds it until export.
  • The payload cannot be read in the Langfuse UI, and the host already stores the file.

Proposed change

Before export, replace each inline payload in span input and output with a short descriptor, keeping the rest of the part (type, MIME type, filename):

{ "type": "media", "mimeType": "application/pdf", "data": "[inline media omitted from trace: application/pdf, 3001128 bytes]" }

Leave data: URIs for the SDK while media upload is enabled, and let hosts opt back into verbatim export.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions