All projects

Personal project, open source (MIT licence)

Mini Data Science Assistant

An agentic assistant for CSV files. Ask in plain English, typed or spoken; a local LLM plans the analysis, chains the right tools and explains the result.

Role
Creator. Architecture, code and agent design
Period
2026
Status
Runs fully locally, code public on GitHub
Stack
  • Python
  • Ollama
  • Gemma
  • FastMCP
  • Pandas
  • Matplotlib
  • Whisper
  • Gradio
How it works
  1. QuestionA CSV file and a question in plain English, typed or spoken (transcribed by Whisper).
  2. AgentA local LLM (Gemma, served by Ollama) plans the analysis and chains up to four tool calls.
  3. MCP serverFastMCP exposes seven data tools built on Pandas. Results move between tools as dataset handles.
  4. AnswerThe model explains the result in plain language, without ever seeing the raw rows.

Gradio interface · runs fully on the local machine, no cloud API

The problem

Exploring a new dataset usually means writing the same Pandas code before you can ask the question you care about. I wanted to ask the question directly, typed or spoken, and get real statistics back, without sending the data to a cloud API.

What I built

  • A multi-step agent loop: a local LLM (Gemma via Ollama) plans the analysis, calls a tool, reads the result and decides the next step, up to four steps per question.
  • An MCP server (FastMCP) with seven data tools: dataset info, statistics, filtering, group-by, missing values, outliers and correlations.
  • In-session dataset handles: results pass from one tool to the next by reference, so the model never sees raw rows.
  • Voice questions transcribed with Whisper, histograms with Matplotlib, and a Gradio interface.
  • Everything runs on the local machine: no cloud API, no data leaving the computer.

My role

I designed the architecture and the agent loop, built the MCP server and its tools, and wrote the interface.