Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Amazon SageMaker Adds Coding Agent Skill for Inference Optimization
AWS released a new skill that turns coding agents into SageMaker inference experts. Engineers can now generate executable code to benchmark models and optimize deployments.
Key points
The aws-ai-ml skill plugs into any agent using the Model Context Protocol (MCP).
Users get quantitative reports on throughput, latency percentiles, and concurrency.
Qwen3-8B showed approximately 44–47 percent higher throughput than Qwen3-1.7B.
Why it matters: Engineers save time by letting agents generate and review optimization code instead of writing it manually.
By OpenSmartRoute editorial · written through the router by writer-small
From AWS machine learning blog - “New agent skill: Amazon SageMaker optimized generative AI inference for your coding agent”
Mona Mona. Image: AWS machine learning blog (original)
Amazon SageMaker adds a new skill for coding agents to optimize inference. Engineers can now generate code to benchmark models and find better hardware. This change helps move projects from idea to production faster. The new tool connects coding assistants directly to AWS infrastructure. It turns general coding agents into experts for AI model performance. You do not need to be an expert in SageMaker to use it. Your agent will handle the complex setup and configuration steps.
New Skill Announcement - Amazon SageMaker introduces the aws-ai-ml skill for coding agents
Amazon SageMaker AI released a new capability for its users today. The company calls this the aws-ai-ml skill. This skill targets coding assistance tools used by software engineers. It gives these agents deep knowledge about inference optimization. Major coding assistants like Kiro, Claude Code, and Codex can now use it. The skill is part of the Agent Toolkit for AWS platform. You install it to unlock specific expertise within your agent.
This announcement addresses a common gap in modern development workflows. Engineers often struggle to bridge the gap between model intent and infrastructure reality. They know they need speed or cost savings but lack the tools to find them. The new skill acts as a specialized plugin for your coding assistant. It plugs into any agent that supports the Model Context Protocol (MCP). This protocol allows agents to interact with external systems safely.
The skill transforms how engineers talk about model performance. Instead of asking vague questions, you describe your business goals. The agent then produces executable code to test those goals. You review the code and run it in your own environment. Nothing happens behind an opaque user interface or hidden menu. Every action is visible as readable text on your screen.
How It Works - The skill plugs into agents via the Model Context Protocol (MCP)
The aws-ai-ml skill works by connecting to agents through a standard protocol. This protocol is called the Model Context Protocol, or MCP for short. It allows coding agents to access AWS resources without needing direct API keys in their memory. The skill configures an AWS MCP Server that acts as a bridge. Your agent talks to this server instead of calling AWS APIs directly.
This design keeps your credentials secure and manageable. You do not need to manage complex IAM roles for the skill itself. The skill generates code that runs under your existing AWS credentials. Your permissions control what the generated code can do. This approach reduces friction during installation and setup. Engineers can adopt the tool without learning a new security model.
The skill integrates seamlessly into existing agent conversations. You start with natural language descriptions of your needs. The agent asks clarifying questions if it lacks necessary details. It then generates Python notebooks or scripts to solve the problem. These outputs are grounded in real benchmark data and measured performance metrics. The agent adapts its behavior based on your specific constraints.
Installation Options - Set up the skill locally or within Amazon SageMaker Studio
Engineers have two main paths to install and use this new skill. You can set it up on your local development machine. Alternatively, you can run it inside an Amazon SageMaker Studio environment. Both methods allow you to go from zero knowledge to a working conversation in about ten minutes. The choice depends on your preferred workflow and security requirements.
For local installation, you use the Agent Toolkit for AWS application. This tool auto-detects your coding agents on your computer. It installs the necessary skills and configures the AWS MCP Server automatically. You need the AWS Command Line Interface version 2.35 or higher. You also need a program called uv installed on your system. These are standard prerequisites for running modern agent tools.
For SageMaker Studio, you use a pre-configured JupyterLab image. This environment is managed by AWS and requires less local setup. You create a new JupyterLab space within your target AWS account. The selected image includes the aws-ai-ml skill and all dependencies. It takes five to ten minutes to boot up on the first run. Subsequent launches are much faster because the environment is cached.
Core Capabilities - Benchmark endpoints and find the right instance types
The core purpose of this skill is to optimize generative AI inference performance. The agent can benchmark existing endpoints deployed on SageMaker AI. It generates Python notebooks that run load tests using specific SDK APIs. These tests measure throughput, latency, and concurrency under real conditions. You do not need to know which API function to call for each metric.
The skill also helps you find the right instance type for any model. This applies whether your model is fine-tuned or a public foundation model. It works with models stored in Amazon S3 or hosted on Hugging Face. The agent evaluates your model against various candidate hardware configurations. It then presents ranked deployment options based on your requirements.
You can compare multiple benchmark runs to see performance deltas. The agent calculates percentage changes across key metrics like throughput and latency. A positive delta indicates an improvement, while a negative one shows degradation. This allows you to make data-driven decisions about configuration changes. You can ask the agent to optimize a model before deploying it to production.
Benchmarking Details - Real load tests measure throughput, latency, and concurrency
Benchmarking in this context means running real traffic against live infrastructure. The skill uses the Workload.synthetic() and start_benchmark() APIs from the SageMaker Python SDK. These functions create synthetic workloads that mimic real user requests. The resulting data is quantitative and based on actual hardware performance. Estimates or theoretical calculations are not used here.
The agent reports specific metrics to give you a complete picture of performance. Throughput measures requests per second and output tokens per second. Latency includes p50, p99, time-to-first-token, and inter-token latency values. Concurrency shows the number of simultaneous requests the system can support. These numbers are measured values from real load on real infrastructure.
Before running any benchmark, the agent confirms that your endpoint is safe to test. It warns you if the benchmark drives significant traffic to a live service. This safety check prevents accidental disruption to production workloads. The agent generates code that includes these confirmation steps automatically. You review the script before executing it in your environment.
The skill handles instance selection logic by comparing models against available hardware. It does not matter where your model currently lives or how you obtained it. You can provide an S3 URI for a custom model or a model ID from JumpStart. For Hugging Face models, the agent checks license terms and asks for tokens if needed.
The agent generates code that evaluates performance across candidate instances. It considers factors like GPU count, memory size, and instance family. The output includes concrete performance metrics for each configuration option. You choose the best option based on your cost and performance requirements. The skill helps you avoid over-provisioning or under-provisioning resources.
If you run multiple benchmarks, the agent can compare them directly. You provide the names of two benchmark jobs to the agent. It computes deltas across key metrics like throughput and latency percentiles. The results show percentage changes where positive means better performance. This gives you a single, interpretable summary of your change's impact.
Why it matters - This skill reduces friction between model selection and production deployment
This skill reduces the friction between selecting a model and deploying it to production. Engineers often spend weeks figuring out infrastructure details manually. The agent automates this process by generating optimized configurations instantly. It saves time on benchmarking, cost analysis, and instance selection tasks.
The skill empowers engineers to make informed decisions without deep AWS expertise. You focus on your business goals like speed or cost savings. The agent handles the technical complexity of finding the right hardware. This leads to faster time-to-market for new AI applications. It also helps control costs by recommending efficient instance types.
Safety is another major benefit of this automated approach. The agent confirms actions before running them to prevent accidental outages. You stay in control by reviewing all generated code and metrics. Nothing happens behind an opaque UI or hidden process. This transparency builds trust in the agentic workflow you are using.
What to do - Install the skill and ask your agent to benchmark or optimize
Start by installing the aws-ai-ml skill through the Agent Toolkit for AWS. Choose either the local installation method or the SageMaker Studio option. Follow the step-by-step instructions provided in the official documentation. Ensure your AWS credentials have permissions to call SageMaker AI APIs. You do not need extra IAM configuration for the skill itself.
Once installed, open your coding agent's chat panel and ask what skills are available. You should see aws-ai-ml listed among the options. Describe your intent in natural language, such as "benchmark my endpoint." The agent will ask clarifying questions if it needs more information. It will then generate executable code to perform the requested task.
Try asking your agent to find the cheapest instance type for a specific model. Or request a comparison between two different benchmark runs. Observe how the agent handles missing information by asking for it rather than guessing. Review the generated Python code carefully before running it in your environment. Delete any resources created during testing to avoid ongoing charges.
Google launched EmbeddingGemma 2, a compact open model that handles text, code, images, video, and audio. It uses a single shared vector space to enable unified search across all media types.