Skip to content

Ask Splitzy: personal spending chatbot (self-hosted LLM) #145

Description

@rghvgrv

Problem Statement

Splitzy users can see totals and balances on the dashboard, but there's no way to ask a free-form question about their own spending — "how much did I spend on groceries this month", "what's my balance", "show grocery expenses over ₹500 last week". Users must mentally filter/sum dashboard data themselves.

Solution

A conversational assistant ("Ask Splitzy") answers questions about the authenticated user's own financial data — spend by category, totals, balances, and ad-hoc filtered queries — using a self-hosted LLM. It never answers about other users' data and refuses anything outside personal-finance scope. Answers stream back token-by-token for responsiveness. Available in both the Angular web app and the Android/Expo app against one shared backend endpoint.

User Stories

  1. As a user, I want to ask "how much did I spend on groceries this month" and get a direct answer, so that I don't have to manually filter expenses.
  2. As a user, I want to ask about my overall total spend and transaction count over a date range, so that I can understand my spending trends.
  3. As a user, I want to ask what my current balance is (who I owe / who owes me), so that I get the same answer the dashboard shows, in conversational form.
  4. As a user, I want to ask open-ended, specific questions ("grocery expenses over $50 last Tuesday") and get a real answer, so that I'm not limited to a fixed set of canned queries.
  5. As a user, I want follow-up questions ("and last month?") to be understood in context, so that the conversation feels natural.
  6. As a user, I want the assistant to only ever discuss my own data, so that I trust it isn't leaking other members' financial information.
  7. As a user, I want the assistant to refuse unrelated requests (general chit-chat, other users' spend), so that it stays trustworthy and on-mission.
  8. As a user, I want to use the assistant equally from the web app and the Android app, so that I have a consistent experience across devices.

Implementation Decisions

  • New module: Chat orchestration service — routes each question through an AI router that classifies it into either a fixed, parameterized query (category totals, overall totals, balance, transaction counts) or a constrained natural-language-to-SQL fallback for open-ended questions; anything off-topic is refused without touching the LLM synthesis step or the database.
  • New module: Chat data-access layer — parameterized query methods scoped server-side to the authenticated user (never LLM-supplied), covering: spend by category, spend by "my share of the split" vs. "what I paid", overall totals, transaction counts, and existing balance/settlement data reused from the current dashboard logic.
  • New module: Constrained query fallback path — for questions the fixed set can't answer, the LLM generates a query which is validated (read-only, single statement, no destructive keywords, mandatory user-scope predicate, row cap) before running against a dedicated read-only, view-restricted database credential. This is a hard security boundary, not just a code-level filter.
  • LLM integration — self-hosted model, no third-party API calls with financial data. Two model roles: a fast router/classifier and a synthesis model that turns retrieved data into a natural-language answer; the fallback query-generation step can reuse either.
  • Conversation contract — multi-turn, stateless on the server: each request carries a bounded window of recent conversation turns; the server does not persist chat history.
  • Response delivery — the final answer streams to the client as it's generated; the routing/data-retrieval steps happen first and are not streamed.
  • Scope enforcement — the authenticated user's identity is always resolved from their auth token server-side and injected into every query path; it is never something the model can set or override. Out-of-scope or unrelated questions get a fixed refusal response and never reach the database or the answer-synthesis step.
  • Client contract — one shared chat API contract consumed by both the web app and the Android app; each client builds its own chat UI (message list, input, streaming display, loading/error/refusal states) against that same contract.
  • Rate limiting — the chat endpoint gets its own, stricter request-rate policy than general API traffic, independent of existing per-user limits.

Testing Decisions

  • Prefer behavior-level tests over implementation-detail tests, consistent with existing backend test practices.
  • Query validator: exhaustively tested against both valid and disallowed inputs (write attempts, multi-statement, missing user-scope, unbounded results) — this is the primary security boundary and must be tested as such.
  • Data-access layer: tested for correctness of category totals, overall totals, transaction counts, and balance figures against known fixture data, mirroring how existing dashboard/balance logic is already covered.
  • Cross-user isolation: an explicit test proving one user's request can never surface another user's rows via either the fixed or fallback query path.
  • Router classification: tested for correct intent selection on representative questions, including out-of-scope refusal cases and malformed/unparseable model output (must fail safe, not fabricate an answer).
  • Client-side: no dedicated automated UI test suite currently exists for comparable features (e.g. OCR flow) in this codebase; manual verification of the chat UI in both apps is acceptable, matching existing project norms.

Out of Scope

  • Answering questions about other users' individual spending (always scoped to self).
  • General-purpose / non-financial chit-chat.
  • Persisted, resumable chat history across sessions or devices.
  • Any write/mutating action performed via the chatbot (it is read-only, informational only).
  • Voice input/output.
  • Multi-currency reasoning beyond what the existing data model already supports.

Further Notes

  • Builds directly on the existing category (ExpenseCategory) and balance/settlement data already computed for the dashboard — no new financial concepts are introduced, only a new, conversational way to query them.
  • Self-hosted model hosting/provisioning is an infrastructure dependency owned by ops and is not part of the application-level slices below.
  • The exact category mapping for informal terms (e.g. "groceries") and the final refusal wording are open product decisions to confirm before/at build time.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or requestneeds-triageNeeds product/eng triage before work starts

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions