Via Jason Liu on LinkedIn

The gap between AI demos and production isn’t about reasoning quality anymore, it’s about context.

Three critical lessons worth stealing:

  1. Domain experts are your secret weapon

• Build custom UIs that optimize for high-quality reviews, not just volume • Collect four outputs from every review: performance metrics, failure modes, suggested improvements, and input-output pairs • Use these to create failure mode datasets you can continuously test against

  1. Prompting beats fine-tuning for vertical AI

• Modern LLMs already have strong domain reasoning, they lack customer-specific context • Skip the brittle prompt engineering and focus on context augmentation instead • Use domain knowledge retrieval at inference time to inject customer-specific definitions and workflows

  1. Trust requires more than good performance

• Monitor production outputs with domain expert reviews (start with humans, not LLM judges) • Define clear sampling strategies as you scale: uncertainty-based, outlier detection, stratified sampling • Have response protocols ready when performance dips below SLAs

An important question to ask is: What’s your system for incorporating domain expertise?

Teams pay too much attention to model selection but neglect the process for continuously baking in expert knowledge. Reliable vertical AI implementations aren’t just about picking the right model, they’re about building flywheels that systematically improve with use.