
Research15 min read
Searching Through a Virtual World
By Nick Mecklenburg
We trained a 4B model to search a synthetic corporate tenant with 91.1% of frontier performance and greater reliability than its 27B teacher.
Research and field notes on evaluations, post-training, and better model behavior.

Research15 min read
By Nick Mecklenburg
We trained a 4B model to search a synthetic corporate tenant with 91.1% of frontier performance and greater reliability than its 27B teacher.

Case Study9 min read
By Jake Kang, Vatsal Bajaj, and Ranjan Goel
Whitney AI took a 27B open-source model and trained it into a red-teamer that outperforms attackers many times its size, 3.1x DeepSeek V4 Flash, with refusal-to-attack down from 81% to 2%.

Research12 min read
By Nick Mecklenburg
No public benchmark for medical billing errors existed, so we built one — synthetic data, start to finish.

Case Study7 min read
By Vatsal Bajaj, Jake Kang, and Nick Mecklenburg
How Whitney AI post-trained a 12B active-parameter generator that writes scam-call transcripts a frontier judge can barely tell from real ones, so Scam AI can build fraud detectors without running thousands of honeypot phone lines.

Case Study7 min read
By Jake Kang
How Whitney AI post-trained a 9B-parameter click-prediction model for CloudCruise's browser-automation builder agent, matching Gemini 3 Flash on accuracy while running 9x faster and 5x cheaper.

Case Study6 min read
By Jake Kang and Nick Mecklenburg
How Whitney AI post-trained a 4B-parameter guardrail model that catches medical-advice violations in Harmonize Health's voice-agent calls before the sentence finishes, matching frontier quality while running 66% faster.

Company3 min read
By Whitney AI Team
Everyone can use a frontier model. Almost no one can shape one. Whitney AI is building the IDE for post-training, so the model you improve is a model you own.