Skip to content

LLMPipeline/Embeddings/Pipeline objects are rebuilt per request instead of being cached singletons #222

Description

@prquinlan

Is there an existing issue for this?

  • I have searched the existing issues

Current Behavior

ai_search() in lettuce/routers/search_routes.py constructs a brand new
LLMPipeline(...).get_rag_assistant() on every call, which itself builds a new
haystack.Pipeline(), ConditionalRouter, PromptBuilder, embedder, and retriever from
scratch. The same pattern exists in lettuce-ui/src/search/__init__.py's _ai_search.

This is the architectural root cause behind the LLM-reload and embedder-reload issues: nothing
expensive in the pipeline graph is hoisted to startup or shared via a cached dependency.

Impact: not a leak in isolation (objects are properly scoped and GC'd), but it guarantees
the expensive resources described in the LLM-reload and embedder-reload issues get recreated on
every request rather than reused.

Suggested fix direction: introduce FastAPI startup-time singletons (e.g. app.state or
Depends + lru_cache) for the LLM client, FastembedTextEmbedder, and the DB engine/session
factory, and reuse them across requests instead of constructing them inside route handlers. This
would resolve the DB session leak, LLM reload, and embedder reload issues together as one piece
of work.

Expected Behavior

No response

Steps To Reproduce

No response

Environment

- OS:
- Other environment details:

I'm part of a Project Team

No response

Anything else?

No response

Are you willing to contribute to resolve this issue?

None

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Fields

    Priority

    None yet

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions