AI / ML engineer

A few things I've built, all runnable right here.

No slides, no screenshots pretending to be demos — six projects below you can actually try before you read a line of code.

Selected work

6 projects, 6 demos
LOG 01

LLM Evaluation & Monitoring Framework

interactive demo

A benchmarking pipeline that runs a curated golden dataset against multiple LLM providers and grades every response on accuracy, hallucination rate, refusal behaviour, and latency (p50/p95) — turning "does this feel right" into a comparable, trackable report across runs and models.

PythonpandasModel adapters
View code ↗

Pick a model to run it against the golden dataset. Scores below are a saved sample run from the framework — re-run the real thing from the repo for live numbers.

Select a model to run the evaluation.
In plain terms: this checks how good and how honest an AI model is before it's trusted in a real product — like a quality inspector, catching when a model makes things up, refuses to answer, or is too slow to be usable.
LOG 02

Evidence-Based Project Health Agent

interactive demo

A project-health reporting agent for enterprise implementations that scores risk deterministically across four dimensions — schedule, execution, milestones, blockers — before an LLM ever sees the data. Gemini's only job is to explain the number and write the narrative; it never decides the rating. Outputs weekly reports and a six-slide executive deck.

PythonGemini APIpython-pptx
View code ↗

Move the sliders the way a project's numbers might move. The RAG rating and explanation update live from a fixed, deterministic formula — no model in the loop.