Skip to content
 
 

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

MediciMess Login Page

MediciMess

Historical Banking Data Engineering & Forensic Analytics Platform

My Role: Project Manager • Data Engineering • Financial Analytics & Anomaly Detection


Project Overview

MediciMess is a team-built data engineering and financial analytics platform that transforms more than 80,000 historically themed banking transactions into validated financial data, branch-level KPIs, anomaly alerts, REST API services, and an interactive operations dashboard.

The project began with a Python implementation of double-entry bookkeeping inspired by the Medici banking dynasty of Renaissance Florence and evolved into an end-to-end data engineering and forensic analytics application.

The completed platform includes:

  • Reusable CSV and JSON ingestion
  • Transaction validation
  • Double-entry accounting controls
  • Financial KPI calculations
  • Branch and monthly aggregation
  • Seven-rule anomaly detection
  • Forensic financial analysis
  • Serving-layer artifacts
  • FastAPI REST services
  • Interactive Dash analytics dashboard
  • Automated testing

The Challenge

The MediciMess dataset contains more than 80,000 historically themed banking transactions spanning 1390–1440.

The data represents banking activity across multiple branches and includes deposits, withdrawals, loans, operating expenses, revenue, and other financial transactions.

Embedded within the dataset is a hidden embezzlement scenario involving approximately 100,000 florins channeled through a fictitious supplier over multiple years.

Our challenge was to build a data pipeline capable of transforming the raw historical transaction data into reliable financial information while also identifying patterns that could indicate suspicious activity.

The project required us to move beyond simply processing transactions. We needed to validate financial data, calculate meaningful business metrics, identify anomalous behavior, expose the resulting information through an API, and make the results accessible through an interactive dashboard.


The Solution

MediciMess evolved into a multi-stage data engineering platform that moves financial information from raw transactions through validation, analytics, anomaly detection, serving, and visualization.

CSV / JSON Transaction Data
            │
            ▼
      Data Ingestion
            │
            ▼
   Transaction Validation
            │
            ▼
   Validated Transactions
            │
            ├───────────────────────┐
            ▼                       ▼
      KPI Analytics          Anomaly Detection
            │                       │
            └───────────┬───────────┘
                        ▼
                  Serving Layer
                        │
                        ▼
                  FastAPI REST API
                        │
                        ▼
                  Dash Dashboard

The architecture separates ingestion, validation, analytics, anomaly detection, serving, and presentation responsibilities so that each stage can be developed, tested, and maintained independently.


My Contributions

Project Manager & Data Engineer

I served as the Project Manager for our three-person development team while also working as a hands-on technical contributor throughout the project.

My role combined project leadership with data engineering, financial analytics, anomaly detection, testing, integration, full-dataset analysis, and technical documentation.

Project Management & Team Leadership

As Project Manager, I coordinated the project across its different development phases while continuing to contribute technically.

My project management responsibilities included:

  • Organizing and prioritizing project work
  • Coordinating work across the three-person team
  • Tracking tasks, dependencies, and development progress
  • Coordinating work between project phases
  • Facilitating team discussions and technical decisions
  • Supporting Git and GitHub collaboration
  • Coordinating integration of team-developed components
  • Reviewing completed functionality
  • Coordinating testing and validation
  • Troubleshooting integration issues
  • Supporting project documentation
  • Coordinating final project readiness and delivery

Data Engineering & Validation

I worked directly with the historical transaction dataset and the pipeline used to prepare the data for downstream analytics.

My work included:

  • Working with the 80,000+ historical transaction dataset
  • Supporting transaction ingestion and validation
  • Verifying financial data before downstream processing
  • Working with CSV and JSON transaction data
  • Reviewing data-quality and pipeline results
  • Running pipeline components against the complete dataset
  • Supporting integration between data-processing phases

The completed ingestion pipeline processed:

Total records: 80,230
Accepted:      80,230
Rejected:           0
Skipped:            0

Financial KPI Analytics

I worked on the financial analytics used to transform validated transactions into branch/month business metrics.

This work included analytics involving:

  • Cash inflows and outflows
  • Net cash movement
  • Closing cash balance
  • Deposits and withdrawals
  • Loan activity
  • Interest
  • Operating expenses
  • Revenue
  • Net income
  • Branch/month aggregation
  • Financial calculations using Python Decimal
  • KPI testing
  • Full-dataset KPI execution

The completed KPI pipeline transformed:

80,230 validated transactions

into:

4,897 branch/month KPI records

Anomaly Detection & Forensic Analytics

I worked on the anomaly-detection and forensic-analysis portion of the project used to identify financial activity requiring additional investigation.

The completed anomaly engine included seven detection rules:

  1. Benford's Law
  2. Vendor Concentration
  3. Duplicate Transactions
  4. Round-Number Clustering
  5. Transaction Frequency Outliers
  6. Structuring / Below-Threshold Activity
  7. New Counterparty High Volume

My work included:

  • Working with anomaly-detection logic
  • Testing detection rules against transaction data
  • Running anomaly analysis against the full dataset
  • Reviewing alert results
  • Investigating unexpected detection behavior
  • Refining analytical logic based on full-dataset results
  • Testing trigger and non-trigger scenarios
  • Validating alert output
  • Supporting integration between KPI and anomaly-processing phases

One significant finding occurred during the Benford's Law analysis.

The initial full-dataset implementation generated:

25,104 Rule A alerts

Analysis showed that the Benford calculation was being applied to groups containing too few transactions to provide a meaningful first-digit distribution.

A minimum sample size of 30 transactions was introduced before applying the calculation.

After the adjustment:

Rule A alerts before: 25,104
Rule A alerts after:     308

This demonstrated an important part of analytical engineering: implementing an algorithm is only the beginning. Its output must also be evaluated against the characteristics of the underlying data and refined when the results reveal limitations in the analytical assumptions.

Testing & Integration

I also worked across testing and integration to help ensure that the individual pipeline components operated correctly together.

My work included:

  • Data-validation testing
  • KPI testing
  • Anomaly-rule testing
  • Trigger and non-trigger testing
  • Full-dataset execution
  • Integration testing
  • Reviewing pipeline outputs
  • Troubleshooting issues identified during testing
  • Supporting integration of team-developed components
  • Reviewing final project functionality

Technical Documentation

I contributed to documenting the project's technical implementation and results, including the data pipeline, KPI analytics, anomaly-detection logic, testing, and full-dataset results.

Working as both Project Manager and Data Engineer allowed me to contribute at two levels: coordinating the team's overall delivery while remaining hands-on with the data, analytics, testing, and technical problem-solving required to build the finished platform.


Data Engineering Pipeline

Data Ingestion

The reusable ingestion layer accepts both CSV and JSON transaction data.

The ingestion process:

  • Reads transaction records
  • Converts IDs to integers
  • Converts dates to Python date objects
  • Converts monetary fields to Decimal
  • Runs shared transaction validation
  • Separates accepted and rejected records
  • Continues processing after malformed records
  • Supports incremental processing
  • Reports potential duplicates without silently removing them

Example:

from ingestion.pipeline import run_pipeline

result = run_pipeline("medici_transactions.csv")

print(f"Accepted: {result.accepted_count}")
print(f"Rejected: {result.rejected_count}")

The complete historical dataset produces:

Total records: 80,230
Accepted:      80,230
Rejected:           0
Skipped:            0

Transaction Validation

Transaction validation helps maintain the integrity of financial data before records are passed downstream for analytics.

The platform builds on double-entry accounting principles, where each financial transaction affects at least two accounts and the sum of debits must equal the sum of credits.

Validation ensures malformed or invalid records can be identified before they affect KPI calculations, anomaly detection, or dashboard reporting.


Financial KPI Analytics

Validated transactions are enriched with time-period information and aggregated by:

branch + monthly period

This transforms individual transaction activity into branch-level financial metrics for downstream reporting and analytics.

KPI Categories

Cash Position

  • Total cash inflows
  • Total cash outflows
  • Net cash movement
  • Closing cash balance

Deposits & Withdrawals

  • Total deposits
  • Total withdrawals
  • Deposit count
  • Withdrawal count
  • Average deposit size
  • Average withdrawal size

Loan Portfolio

  • Loans issued
  • Loans repaid
  • Interest earned
  • Loan portfolio balance
  • Interest yield

Operating Expenses

  • Total operating expenses
  • Expenses by category
  • Expense per transaction
  • Top payees by expense

Revenue & Profitability

  • Exchange fee revenue
  • Interest income
  • Trading revenue
  • Total revenue
  • Net income
  • Net income margin

Financial calculations use Python Decimal values rather than floating-point values to maintain appropriate precision for monetary calculations.


Full Dataset KPI Results

The KPI pipeline was executed against the complete validated dataset.

Source transactions:    80,230
Accepted transactions:  80,230
Rejected transactions:       0
KPI records generated:   4,897

The reduction from 80,230 transactions to 4,897 KPI records occurs because transaction data is aggregated into branch/month reporting periods.


Anomaly Detection

The platform includes a seven-rule anomaly detection engine designed to identify financial activity that may require additional investigation.

Rather than relying on a single technique, the engine examines multiple types of potentially unusual financial behavior.

Rule Detection Method Purpose
A Benford's Law Identify unusual first-digit distributions
B Vendor Concentration Detect unusually concentrated spending
C Duplicate Transactions Identify potential duplicate activity
D Round-Number Clustering Detect concentrations of round-number transactions
E Transaction Frequency Outliers Identify sudden increases in transaction frequency
F Structuring Detect repeated below-threshold activity
G New Counterparty High Volume Identify unusually high activity from new counterparties

Benford's Law

Rule A analyzes first-digit distributions by:

branch + transaction type + monthly period

Mean Absolute Deviation is calculated between the observed first-digit distribution and the expected Benford distribution.

Configuration:

MAD threshold:        0.015
Minimum sample size: 30 transactions

Full-dataset analysis initially generated 25,104 Rule A alerts.

Inspection showed that analysis was being performed on groups containing too few transactions to provide a meaningful first-digit distribution.

After introducing the 30-transaction minimum sample requirement:

Rule A alerts before: 25,104
Rule A alerts after:     308

Vendor Concentration

Rule B analyzes operating expenses by:

branch + period + expense category + counterparty

The analysis identifies situations where a counterparty represents an unusually large share of spending within an expense category.


Duplicate Transactions

Rule C identifies potential duplicate transactions using:

  • Transaction type
  • Counterparty
  • Debit amount
  • Transaction date

Matching transactions occurring within three calendar days are flagged for additional review.


Round-Number Clustering

Rule D evaluates operating expenses for unusual concentrations of round-number transactions.

This can help identify patterns that may warrant further financial review.


Transaction Frequency Outliers

Rule E analyzes monthly transaction frequency for:

branch + counterparty + transaction type

This helps identify sudden increases in activity involving a particular counterparty.


Structuring / Below-Threshold Activity

Rule F identifies repeated transactions below a defined individual threshold that collectively exceed an aggregate threshold during the reporting period.

Transactions are grouped by:

branch + period + counterparty + transaction type

New Counterparty High Volume

Rule G identifies new counterparties with unusually high transaction activity during their first active reporting period at a branch.


Full Dataset Anomaly Results

The anomaly engine was executed against all 80,230 accepted transactions.

Rule Detection Alerts
A Benford's Law 308
B Vendor Concentration 9,011
C Duplicate Transactions 0
D Round-Number Clustering 25
E Transaction Frequency Outliers 1,263
F Structuring 0
G New Counterparty High Volume 176
Total 10,783

A zero-alert result does not indicate that a rule failed. Trigger and non-trigger tests verify that the rules generate alerts when their defined conditions are present.

The alerts identify activity for additional investigation rather than automatically classifying transactions as fraudulent.


Forensic Analysis

The historical transaction dataset contains a hidden embezzlement scenario within the Florence branch's operating expenses during 1420–1424.

The scenario involves approximately 100,000 florins channeled through a fictitious supplier over five years.

The platform provides analytical techniques that can be used to investigate suspicious activity, including:

  • Benford's Law
  • Vendor concentration analysis
  • Transaction frequency analysis
  • Counterparty analysis
  • Round-number analysis
  • Transaction-pattern analysis

This transformed the original banking simulation into a forensic financial data-analysis exercise.


Serving Layer

The serving layer consumes completed KPI/detail records and anomaly-alert records.

It validates, partitions, and writes serving artifacts without recalculating financial metrics or anomaly rules.

This separation allows downstream applications to consume prepared analytical outputs without duplicating business logic.

Generate artifacts from the complete historical dataset:

python3 generate_serving_artifacts.py medici_transactions.csv \
    --output serving_outputs

FastAPI REST API

The platform exposes analytical data through a read-only FastAPI REST API.

The API consumes serving artifacts and the validated transaction ledger without recalculating KPI or anomaly logic.

Available endpoints include:

GET /health
GET /api/branches
GET /api/kpis
GET /api/transactions
GET /api/cashflow
GET /api/loans
GET /api/expenses
GET /api/alerts

Start the development server:

python3 -m pip install -r requirements.txt
uvicorn api.app:app --reload

Interactive API documentation is available locally at:

http://127.0.0.1:8000/docs

Branch Operations Dashboard

The Dash Branch Operations Dashboard sits on top of the FastAPI service and provides an interactive interface for exploring banking operations.

The dashboard includes:

  • Global branch and period controls
  • Financial KPI comparisons
  • Cash-flow visualizations
  • Loan analysis
  • Expense analysis
  • Bills-of-exchange activity
  • Anomaly review
  • Searchable transaction ledger
  • Sorting and pagination

Start Dash after starting FastAPI:

MEDICIMESS_API_URL=http://127.0.0.1:8000 \
python3 -m dashboard.app

Open:

http://127.0.0.1:8050

Team

MediciMess_Charlie was developed collaboratively by a three-person team.

Team Member Role / Primary Contribution
Leigh Project Manager • Data Engineering • Financial Analytics & Anomaly Detection
Hakeem Team Developer
Vijay Team Developer

Development, testing, integration, and delivery required collaboration across the team.


Tech Stack

Languages & Data Processing

  • Python
  • CSV
  • JSON
  • Python Decimal

Data Engineering & Analytics

  • Data ingestion
  • Data validation
  • Data transformation
  • Financial KPI calculations
  • Branch/month aggregation
  • Anomaly detection
  • Forensic data analysis

API & Application

  • FastAPI
  • Uvicorn
  • Dash

Development & Testing

  • Git
  • GitHub
  • Pytest
  • REST APIs
  • Automated testing

Project Structure

Key components include:

ingestion/
    csv_ingestion.py
    json_ingestion.py
    pipeline.py

analytics/
    account_types.py
    kpis.py
    alerts.py

api/
    app.py

dashboard/
    app.py

tests/
    test_csv_ingestion.py
    test_json_ingestion.py
    test_validation.py
    test_kpis.py
    test_alerts.py

This separation of ingestion, analytics, API, dashboard, and testing components supports a modular application architecture.


Running the Project

Requirements

  • Python 3
  • Project dependencies from requirements.txt

Clone your fork:

git clone https://github.com/LDurham1213/MediciMess_Charlie.git
cd MediciMess_Charlie

Install dependencies:

python3 -m pip install -r requirements.txt

Run the Data Pipeline

from ingestion.pipeline import run_pipeline

result = run_pipeline("medici_transactions.csv")

print(f"Accepted: {result.accepted_count}")
print(f"Rejected: {result.rejected_count}")

Generate Serving Artifacts

python3 generate_serving_artifacts.py medici_transactions.csv \
    --output serving_outputs

Start FastAPI

uvicorn api.app:app --reload

API documentation:

http://127.0.0.1:8000/docs

Start the Dash Dashboard

After starting FastAPI:

MEDICIMESS_API_URL=http://127.0.0.1:8000 \
python3 -m dashboard.app

Open:

http://127.0.0.1:8050

Project Documentation

Additional technical documentation is available throughout the repository:

  • docs/PIPELINE_DOCUMENTATION.md — pipeline architecture and implementation
  • docs/USER_GUIDE.md — application usage
  • docs/DEPLOYMENT_RUNBOOK.md — deployment and operational instructions
  • docs/API_REFERENCE.md — REST API documentation
  • notebooks/kpi_anomaly_rules_demo.ipynb — KPI and anomaly examples
  • DATA_CONTRACTS.md — pipeline handoff structures
  • DATA_PIPELINE_SPEC.md — pipeline requirements and architecture
  • PHASE5_SERVING_LAYER.md — serving-layer implementation
  • PHASE6_README.md — API and dashboard implementation

Project Evolution

MediciMess began as a Python implementation of double-entry bookkeeping.

The project was expanded into a much larger data engineering challenge using more than 80,000 historically themed financial transactions.

The completed platform evolved through several layers:

Double-Entry Accounting
        ↓
Historical Transaction Data
        ↓
Reusable Data Ingestion
        ↓
Transaction Validation
        ↓
Financial KPI Analytics
        ↓
Anomaly Detection
        ↓
Forensic Analysis
        ↓
Serving Layer
        ↓
FastAPI
        ↓
Dash Dashboard

This progression transformed the original accounting exercise into an end-to-end financial data engineering and analytics platform.


Key Takeaways

MediciMess demonstrated how raw transactional data can be transformed into reliable, decision-ready information through a layered data engineering architecture.

The project required more than simply loading and displaying data. Financial records had to be validated, aggregated into meaningful business metrics, evaluated for anomalous behavior, prepared for downstream consumption, exposed through an API, and presented through an interactive application.

As Project Manager and a hands-on technical contributor, this project allowed me to combine:

  • Project leadership
  • Team coordination
  • Data engineering
  • Data validation
  • Financial analytics
  • KPI development
  • Anomaly detection
  • Forensic data analysis
  • Python development
  • REST API architecture
  • Automated testing
  • Git/GitHub collaboration
  • Application integration
  • Technical documentation

One of the most valuable lessons from the project was that successful anomaly detection is not simply about producing alerts. Detection logic must be tested against the characteristics of the underlying data, results must be interpreted in context, and analytical assumptions must be refined when the data shows that they are not producing meaningful results.


Project Status

Completed Team Project

MediciMess_Charlie was completed as a collaborative data engineering and analytics project.

This fork is maintained as part of my professional portfolio to document both the team's completed application and my contributions as Project Manager and Data Engineer.


License

This project is licensed under the MIT License.

See the LICENSE file for additional information.

About

Team-built data engineering and analytics platform for historical banking data, featuring ETL pipelines, financial KPIs, anomaly detection, FastAPI, and Dash.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages