Optimizing Agent Behavior with LangSmith
Podium, a communication platform for small businesses, leveraged LangSmith to optimize their AI Employee agentic application and significantly reduce the time engineers spent on troubleshooting. Initially, the team utilized LangChain for single-turn interactions, but as their agentic use cases expanded, they needed a robust tool for monitoring LLM calls and interactions. LangSmith provided the visibility required to manage a growing set of complex scenarios.
Podium established a comprehensive testing lifecycle incorporating dataset curation, offline evaluation, user-provided feedback, and online monitoring. This involved creating baseline datasets, conducting initial tests, and continuously updating the dataset with new edge cases. The team utilized techniques like model distillation to refine their models, improving performance and reducing model size.
Specifically, Podium addressed an issue where the AI Employee repeatedly said ‘goodbye’ even after a natural conversation end. By curating a dataset in LangSmith with diverse conversation scenarios, they were able to train the model to recognize when a conversation should conclude. This resulted in a 7.5% improvement in F1 scores, increasing from 91.7% to 98.6% – exceeding their quality threshold.
Furthermore, the Technical Product Specialists (TPS) at Podium gained access to LangSmith, enabling them to quickly identify and resolve customer-reported issues in real-time. This allowed them to determine whether issues stemmed from application bugs, incomplete context, misaligned instructions, or LLM problems. This improved operational efficiency and facilitated faster product improvements.



