LLM Evaluation & Monitoring Framework
interactive demoA benchmarking pipeline that runs a curated golden dataset against multiple LLM providers and grades every response on accuracy, hallucination rate, refusal behaviour, and latency (p50/p95) — turning "does this feel right" into a comparable, trackable report across runs and models.
Pick a model to run it against the golden dataset. Scores below are a saved sample run from the framework — re-run the real thing from the repo for live numbers.