← All Case Studies
AI Infrastructure · MLOps

Self-Hosted LLM Evaluation & Regression Platform

Internal Product · ChuksForge
Self-Hosted LLM Evaluation & Regression Platform
The Challenge

Prompt and model changes across our own AI products were shipping on manual review, the exact gap we later wrote about publicly. There was no self-hosted, CI-integrated way to compare a change against an explicitly approved baseline and block a regression before it merged, and every eval platform evaluated offered that only behind an enterprise-only self-hosting tier.

The Solution

A self-hosted evaluation and regression platform: LLM-as-judge and lexical scoring, explicit baseline comparison instead of a drifting rolling average, and severity-aware CI exit codes that actually gate a merge. The detection and regression engine was open-sourced separately (Apache-2.0) so the core logic is auditable independent of the hosted platform, and the platform itself now runs the same evals-in-CI discipline behind features like EvalCI.

Tech Stack
PythonFastAPICelery BeatPostgreSQLLLM-as-JudgeApache-2.0 Core Engine
CI-gated
Every prompt/model change
Apache-2.0
Open-sourced detection engine
Self-hosted
No enterprise-tier lock-in

Have a similar challenge?

Book a free 30-minute architecture call and we'll tell you honestly whether and how we can help.

Book Free Discovery Call →