Skip to content

About

An AI-powered tool that generates plain-language metadata descriptions for government datasets on data.wa.gov

Resources

Stars

0 stars

Watchers

1 watching

Forks

Repository files navigation

AI Metadata Improvement Tool

A Generative AI tool designed to assist data publishers in improving the quality, consistency, and accessibility of metadata on the Washington State Open Data Portal (data.wa.gov).

Project Overview

This project was developed as part of the MSIM Capstone program in partnership with Washington Technology Solutions and the State of Washington Open Data Program.

High-value datasets regarding state licensing, transportation, healthcare, and fiscal matters are hosted on the Socrata platform. However, metadata often falls short of completeness or fails to use "plain language," making it difficult for the public to utilize these resources. This tool addresses that gap by using AI to generate compliant, descriptive metadata automatically.

Goals and Objectives

  • Automate Metadata Generation: Utilize LLMs to analyze dataset samples and schemas to suggest titles, descriptions, and column definitions automatically
  • Enhance Accessibility & Consistency: Enforce Plain Language standards (expanding acronyms, simplifying jargon) and ensure consistency with U.S. open data standards
  • Enable User Iteration: Create a "Human-in-the-Loop" workflow that lets publishers accept, reject, or regenerate suggestions with specific instructions
  • Cost & Performance Optimization: Generate high-quality descriptions quickly and cost-effectively for sustainable public-sector use
  • Platform Independence: Free, open-access tool deployable without reliance on ongoing subscriptions

Features

  • Flexible Data Import: Connect directly to Socrata (data.wa.gov) via dataset ID or upload local CSV files.
  • Smart Column Analysis: Automatic detection of data types (numeric, categorical, text) with statistical summaries.
  • AI-Powered Metadata: Real-time generation of titles, descriptions, row labels, and temporal metadata.
  • Universal LLM Support: Compatible with any OpenAI-compatible API (Azure, Databricks, Ollama, HuggingFace, etc.).
  • Human-in-the-Loop Workflow: Review, edit, and iterate on AI suggestions with custom instructions and streaming responses.
  • Direct Socrata Export: Push metadata updates back to the portal via OAuth or API Key authentication.
  • Full Customization: Modify AI prompts, system personas, and regeneration styles to match specific guidelines.
  • Cost Efficiency: Built-in token usage tracking and cost estimation for sustainable public-sector use.

Local Development

Prerequisites

Installation

  1. Install frontend dependencies:

    npm install
  2. Install backend dependencies:

    cd backend
    python3 -m venv venv
    source venv/bin/activate  # On Windows: venv\Scripts\activate
    pip install -r requirements.txt
    cd ..

Environment Setup

  1. Copy the example environment files:

    cp .env.example .env
    cp backend/.env.example backend/.env
  2. Configure the backend .env (in the backend/ directory):

    • Set SOCRATA_APP_TOKEN (get one from data.wa.gov)
    • (Optional) Set LLM_ENDPOINT, LLM_API_KEY, and LLM_MODEL to pre-configure the AI

Running the App

You can start both the frontend and backend with a single command:

npm run dev:all

Alternatively, you can start them in separate terminals:

Terminal 1: Backend

# Activate the virtual environment if you haven't already
source backend/venv/bin/activate
python -m backend.main

Terminal 2: Frontend

npm run dev

Socrata OAuth Setup (Sign in with data.wa.gov)

OAuth login allows users to authenticate with their own portal credentials. Socrata requires HTTPS for OAuth callbacks, so you must use a tunnel for local development.

  1. Expose your local backend: Use ngrok or similar to create an HTTPS tunnel to port 8000:
    ngrok http 8000
  2. Register an App Token: On data.wa.gov, create a new app token and set the Callback Prefix to your ngrok URL: https://your-tunnel-id.ngrok-free.app/api/auth/socrata/
  3. Update backend/.env:
    SOCRATA_SECRET_TOKEN=your-secret-token
    # Override the redirect URI explicitly because the backend (port 8000) and
    # the frontend dev server (port 5173) live on different origins in local dev.
    SOCRATA_OAUTH_REDIRECT_URI=https://your-tunnel-id.ngrok-free.app/api/auth/socrata/callback
    FRONTEND_URL=http://localhost:5173

Note: The ngrok URL may change each time you restart the tunnel (free tier). You will need to update the Callback Prefix and SOCRATA_OAUTH_REDIRECT_URI accordingly.

In production (Databricks Apps), the frontend and backend share an origin, so SOCRATA_OAUTH_REDIRECT_URI is derived automatically from FRONTEND_URL and does not need to be set.

Usage

  1. Configure LLM Provider: Enter your API base URL, API key, and model name (or pre-configure via environment variables).
  2. Import Data:
    • Enter a Socrata dataset ID (e.g., 6fex-3r7d) to import from data.wa.gov, or
    • Paste a full dataset URL — the tool extracts both the dataset ID and the portal, switching to that portal automatically (e.g. a data.cityofnewyork.us URL), or
    • Upload a local CSV file
    • To target a different portal directly, set it under Settings → Socrata Portal
  3. Review Results: View the generated dataset title, description, row label, category, tags, and individual column descriptions with real-time streaming
  4. Iterate:
    • Edit descriptions inline
    • Regenerate with "More Concise", "More Detailed", or custom instructions
  5. Export: Push updated metadata back to data.wa.gov (requires OAuth or API Key authentication)

Customizing Prompts

The system relies on prompt templates to steer the AI's behavior when generating metadata. All default prompt templates are stored as plain .md files in the prompts/ directory at the root of the project.

To update or customize the AI prompts (e.g., system personas, formatting instructions, or dataset criteria):

  1. Navigate to the prompts/ directory.
  2. Open the relevant .md file (such as system.md, dataset.md, or column.md).
  3. Make your edits and save the file.
  4. The changes will be applied automatically if you are running the local development server.

Note: Ensure you do not remove or modify the untrusted-data fence tokens (<<<UNTRUSTED_DATA>>> and <<<END_UNTRUSTED_DATA>>>) present in the prompt templates, as these securely isolate dataset content and prevent prompt injection attacks.

Deployment

Quick Start for Testers (Manual Setup)

If you want to quickly test the tool in your own Databricks workspace without setting up GitHub Actions, follow these steps:

  1. Create a Git Folder: Go to Workspace, and select Create → Git folder.
  2. Point to the release-databricks branch:
  3. Configure Secrets:
    • Inside the Git Folder, create a file named .env.databricks in the root directory.
    • Copy the content from .env.databricks.example into it and fill in your keys (OpenAI, Socrata, etc.).
    • (Optional) You can safely remove DATABRICKS_APP_NAME and DATABRICKS_WORKSPACE_PATH — they're only used by the automated deploy.
    • Leave FRONTEND_URL as a placeholder for now; you'll set it in step 4 once you know the app's URL.
  4. Create the App and Set FRONTEND_URL:
    • Go to Compute → Apps → Create app.
    • Choose Custom app, name it, and click Create. (The Source code path can't be set at creation time.)
    • Once the app exists, copy its URL (e.g., https://your-app-id.databricksapps.com) from the app page and set FRONTEND_URL in .env.databricks to that value. The Socrata OAuth redirect URI is derived from this automatically.
  5. Deploy the App:
    • On the app page, click Deploy — Databricks will prompt you for the Source code path; point it at your Git Folder.
    • The app will build and start, picking up the values you set in .env.databricks.

For automated production deployments, see DEPLOYMENT.md.

OpenAI-Compatible API Support

This tool works with any LLM provider that implements the OpenAI chat completion API. The backend uses the OpenAI Python SDK with a configurable base_url, so any service exposing a compatible /v1/chat/completions endpoint will work out of the box.

Supported Providers

Provider Type
Databricks Cloud
OpenAI Cloud
Microsoft Azure Foundry Cloud
Ollama Local / Cloud
LM Studio Local
HuggingFace Cloud
Any OpenAI-compatible server Either

Recommended Starter Models

Model Notes
gpt-5-mini Fast and cost-effective
gpt-5-nano Lightest OpenAI option
Qwen3-4B-Instruct-2507 Small, runs well locally
Qwen/Qwen3-8B Strong open-weight model
mistralai/Ministral-3-8B-Instruct-2512 Compact Mistral model
mistralai/Ministral-3-14B-Instruct-2512 Higher capacity Mistral model

Team

Name Role
Wynter Lin AI & Cloud Computing
Danny Yue UI/UX & Machine Learning
Felix Zhao DS & Backend Development
Julia Zhu BI & Data Visualization

Acknowledgments

  • Washington State Open Data Program
  • Cathi Greenwood, Open Data Program Manager
  • Kathleen Sullivan, Open Data Literacy Consultant

License

This project is licensed under the Apache License 2.0. It is open-access and free to use, designed for replication by other government data portals using Tyler Technologies Data & Insights.

About

An AI-powered tool that generates plain-language metadata descriptions for government datasets on data.wa.gov

Resources

Stars

0 stars

Watchers

1 watching

Forks

Contributors

Languages