Skills let you encode domain-specific procedures as reusable, portable instructions for agents, but a fluent answer doesn't prove the agent picked the right skill or followed it. Learn how to measure skill selection and instruction following with Strands Evals and Amazon Bedrock AgentCore Evaluations.
Read the original at AWS machine learning blog: Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore


![[AINews] Claude Opus 5.5, the new default model for AINews — and everybody cuts prices 40-50%](https://substackcdn.com/image/fetch/$s_!tLrK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F__ss-rehost__tw-video-preview-13_2102432417352929280.jpg)
