Sit with the model, debug why it's behaving unexpectedly, design synthetic examples that teach the desired behavior, measure with evals, repeat.
Launch, observe real user behavior patterns, shift training distribution to match how users actually use it (not how you predicted).
Send us yours and we’ll draw it. $1 a transcript.