Personal project, open source (MIT licence)
Mini Data Science Assistant
An agentic assistant for CSV files. Ask in plain English, typed or spoken; a local LLM plans the analysis, chains the right tools and explains the result.
- Role
- Creator. Architecture, code and agent design
- Period
- 2026
- Status
- Runs fully locally, code public on GitHub
- Stack
- Python
- Ollama
- Gemma
- FastMCP
- Pandas
- Matplotlib
- Whisper
- Gradio
- QuestionA CSV file and a question in plain English, typed or spoken (transcribed by Whisper).
- AgentA local LLM (Gemma, served by Ollama) plans the analysis and chains up to four tool calls.
- MCP serverFastMCP exposes seven data tools built on Pandas. Results move between tools as dataset handles.
- AnswerThe model explains the result in plain language, without ever seeing the raw rows.
Gradio interface · runs fully on the local machine, no cloud API
The problem
Exploring a new dataset usually means writing the same Pandas code before you can ask the question you care about. I wanted to ask the question directly, typed or spoken, and get real statistics back, without sending the data to a cloud API.
What I built
- A multi-step agent loop: a local LLM (Gemma via Ollama) plans the analysis, calls a tool, reads the result and decides the next step, up to four steps per question.
- An MCP server (FastMCP) with seven data tools: dataset info, statistics, filtering, group-by, missing values, outliers and correlations.
- In-session dataset handles: results pass from one tool to the next by reference, so the model never sees raw rows.
- Voice questions transcribed with Whisper, histograms with Matplotlib, and a Gradio interface.
- Everything runs on the local machine: no cloud API, no data leaving the computer.
My role
I designed the architecture and the agent loop, built the MCP server and its tools, and wrote the interface.