A Generative AI tool designed to assist data publishers in improving the quality, consistency, and accessibility of metadata on the Washington State Open Data Portal (data.wa.gov).
This project was developed as part of the MSIM Capstone program in partnership with Washington Technology Solutions and the State of Washington Open Data Program.
High-value datasets regarding state licensing, transportation, healthcare, and fiscal matters are hosted on the Socrata platform. However, metadata often falls short of completeness or fails to use "plain language," making it difficult for the public to utilize these resources. This tool addresses that gap by using AI to generate compliant, descriptive metadata automatically.
- Automate Metadata Generation: Utilize LLMs to analyze dataset samples and schemas to suggest titles, descriptions, and column definitions automatically
- Enhance Accessibility & Consistency: Enforce Plain Language standards (expanding acronyms, simplifying jargon) and ensure consistency with U.S. open data standards
- Enable User Iteration: Create a "Human-in-the-Loop" workflow that lets publishers accept, reject, or regenerate suggestions with specific instructions
- Cost & Performance Optimization: Generate high-quality descriptions quickly and cost-effectively for sustainable public-sector use
- Platform Independence: Free, open-access tool deployable without reliance on ongoing subscriptions
- Flexible Data Import: Connect directly to Socrata (data.wa.gov) via dataset ID or upload local CSV files.
- Smart Column Analysis: Automatic detection of data types (numeric, categorical, text) with statistical summaries.
- AI-Powered Metadata: Real-time generation of titles, descriptions, row labels, and temporal metadata.
- Universal LLM Support: Compatible with any OpenAI-compatible API (Azure, Databricks, Ollama, HuggingFace, etc.).
- Human-in-the-Loop Workflow: Review, edit, and iterate on AI suggestions with custom instructions and streaming responses.
- Direct Socrata Export: Push metadata updates back to the portal via OAuth or API Key authentication.
- Full Customization: Modify AI prompts, system personas, and regeneration styles to match specific guidelines.
- Cost Efficiency: Built-in token usage tracking and cost estimation for sustainable public-sector use.
- Node.js 24+ (for frontend build)
- Python 3.12+ (for backend)
- An LLM provider — any OpenAI-compatible API (e.g., Ollama, LM Studio, HuggingFace, or OpenAI)
- A Socrata App Token — required for fetching metadata from data.wa.gov. See Developer Settings.
-
Install frontend dependencies:
npm install
-
Install backend dependencies:
cd backend python3 -m venv venv source venv/bin/activate # On Windows: venv\Scripts\activate pip install -r requirements.txt cd ..
-
Copy the example environment files:
cp .env.example .env cp backend/.env.example backend/.env
-
Configure the backend
.env(in thebackend/directory):- Set
SOCRATA_APP_TOKEN(get one from data.wa.gov) - (Optional) Set
LLM_ENDPOINT,LLM_API_KEY, andLLM_MODELto pre-configure the AI
- Set
You can start both the frontend and backend with a single command:
npm run dev:allAlternatively, you can start them in separate terminals:
Terminal 1: Backend
# Activate the virtual environment if you haven't already
source backend/venv/bin/activate
python -m backend.mainTerminal 2: Frontend
npm run dev- Frontend: http://localhost:5173
- Backend: http://localhost:8000
OAuth login allows users to authenticate with their own portal credentials. Socrata requires HTTPS for OAuth callbacks, so you must use a tunnel for local development.
- Expose your local backend: Use ngrok or similar to create an HTTPS tunnel to port 8000:
ngrok http 8000
- Register an App Token: On data.wa.gov, create a new app token and set the Callback Prefix to your ngrok URL:
https://your-tunnel-id.ngrok-free.app/api/auth/socrata/ - Update
backend/.env:SOCRATA_SECRET_TOKEN=your-secret-token # Override the redirect URI explicitly because the backend (port 8000) and # the frontend dev server (port 5173) live on different origins in local dev. SOCRATA_OAUTH_REDIRECT_URI=https://your-tunnel-id.ngrok-free.app/api/auth/socrata/callback FRONTEND_URL=http://localhost:5173
Note: The ngrok URL may change each time you restart the tunnel (free tier). You will need to update the Callback Prefix and
SOCRATA_OAUTH_REDIRECT_URIaccordingly.In production (Databricks Apps), the frontend and backend share an origin, so
SOCRATA_OAUTH_REDIRECT_URIis derived automatically fromFRONTEND_URLand does not need to be set.
- Configure LLM Provider: Enter your API base URL, API key, and model name (or pre-configure via environment variables).
- Import Data:
- Enter a Socrata dataset ID (e.g.,
6fex-3r7d) to import from data.wa.gov, or - Paste a full dataset URL — the tool extracts both the dataset ID and the
portal, switching to that portal automatically (e.g. a
data.cityofnewyork.usURL), or - Upload a local CSV file
- To target a different portal directly, set it under Settings → Socrata Portal
- Enter a Socrata dataset ID (e.g.,
- Review Results: View the generated dataset title, description, row label, category, tags, and individual column descriptions with real-time streaming
- Iterate:
- Edit descriptions inline
- Regenerate with "More Concise", "More Detailed", or custom instructions
- Export: Push updated metadata back to data.wa.gov (requires OAuth or API Key authentication)
The system relies on prompt templates to steer the AI's behavior when generating metadata. All default prompt templates are stored as plain .md files in the prompts/ directory at the root of the project.
To update or customize the AI prompts (e.g., system personas, formatting instructions, or dataset criteria):
- Navigate to the
prompts/directory. - Open the relevant
.mdfile (such assystem.md,dataset.md, orcolumn.md). - Make your edits and save the file.
- The changes will be applied automatically if you are running the local development server.
Note: Ensure you do not remove or modify the untrusted-data fence tokens (
<<<UNTRUSTED_DATA>>>and<<<END_UNTRUSTED_DATA>>>) present in the prompt templates, as these securely isolate dataset content and prevent prompt injection attacks.
If you want to quickly test the tool in your own Databricks workspace without setting up GitHub Actions, follow these steps:
- Create a Git Folder: Go to Workspace, and select Create → Git folder.
- Point to the
release-databricksbranch:- URL: https://github.com/HuskyDevClub/AI-Metadata-Improvement-Tool.git
- Branch:
release-databricks.
- Configure Secrets:
- Inside the Git Folder, create a file named
.env.databricksin the root directory. - Copy the content from
.env.databricks.exampleinto it and fill in your keys (OpenAI, Socrata, etc.). - (Optional) You can safely remove
DATABRICKS_APP_NAMEandDATABRICKS_WORKSPACE_PATH— they're only used by the automated deploy. - Leave
FRONTEND_URLas a placeholder for now; you'll set it in step 4 once you know the app's URL.
- Inside the Git Folder, create a file named
- Create the App and Set
FRONTEND_URL:- Go to Compute → Apps → Create app.
- Choose Custom app, name it, and click Create. (The Source code path can't be set at creation time.)
- Once the app exists, copy its URL (e.g.,
https://your-app-id.databricksapps.com) from the app page and setFRONTEND_URLin.env.databricksto that value. The Socrata OAuth redirect URI is derived from this automatically.
- Deploy the App:
- On the app page, click Deploy — Databricks will prompt you for the Source code path; point it at your Git Folder.
- The app will build and start, picking up the values you set in
.env.databricks.
For automated production deployments, see DEPLOYMENT.md.
This tool works with any LLM provider that implements the OpenAI chat completion API. The backend uses the OpenAI Python SDK with a configurable base_url, so any service exposing a compatible /v1/chat/completions endpoint will work out of the box.
| Provider | Type |
|---|---|
| Databricks | Cloud |
| OpenAI | Cloud |
| Microsoft Azure Foundry | Cloud |
| Ollama | Local / Cloud |
| LM Studio | Local |
| HuggingFace | Cloud |
| Any OpenAI-compatible server | Either |
| Model | Notes |
|---|---|
gpt-5-mini |
Fast and cost-effective |
gpt-5-nano |
Lightest OpenAI option |
Qwen3-4B-Instruct-2507 |
Small, runs well locally |
Qwen/Qwen3-8B |
Strong open-weight model |
mistralai/Ministral-3-8B-Instruct-2512 |
Compact Mistral model |
mistralai/Ministral-3-14B-Instruct-2512 |
Higher capacity Mistral model |
| Name | Role |
|---|---|
| Wynter Lin | AI & Cloud Computing |
| Danny Yue | UI/UX & Machine Learning |
| Felix Zhao | DS & Backend Development |
| Julia Zhu | BI & Data Visualization |
- Washington State Open Data Program
- Cathi Greenwood, Open Data Program Manager
- Kathleen Sullivan, Open Data Literacy Consultant
This project is licensed under the Apache License 2.0. It is open-access and free to use, designed for replication by other government data portals using Tyler Technologies Data & Insights.