Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Together Link puts frontier open models in your existing coding agents - OpenSmartRoute
Together Link puts frontier open models in your existing coding agents
Together AI launched Together Link to connect popular coding agents with high-quality open models. This setup cuts engineering spend by over 50% while keeping the same workflow.
Key points
Together Link reduces monthly costs by over 50% for coding agents.
The tool supports over 40 models chosen for production use.
Kimi K3 and GLM 5.3 now handle hard coding tasks at lower prices.
Together AI serves the largest share of OpenRouter tokens for top open models.
Why it matters: Switching to cheaper open models saves engineering teams tens of thousands of dollars monthly without changing their agent setup.
By OpenSmartRoute editorial · written through the router by writer-small
From Together AI blog - “Together Link: open models in the harness you already use. Start with one command today.”
Together AI launched a new tool called Together Link on October 5, 2026. This tool connects popular coding agents to high-quality open models. It lets teams use these models without changing their current workflow. The announcement highlights that engineering spend can drop by over 50%. Engineers do not need to learn new settings or logins. They can switch back to native closed models with one command if needed.
This product aims to bring frontier open models into the harness your team already uses. It replaces expensive closed models for many coding tasks. The goal is to keep the same agent experience while lowering costs. Teams can run these models on serverless infrastructure immediately. No separate contracts or billing systems are required for this setup.
How coding agents currently spend money on closed models
Coding agents have become common tools for engineering teams worldwide. Every task given to them runs on premium models often. These tasks range from one-line fixes to full code rewrites. At a large scale, this usage gets very expensive quickly. Engineering organizations spend tens of thousands of dollars a month. Some orgs spend millions of dollars monthly on these closed models.
Most coding agents rely on closed models for their core functionality. Closed models are proprietary software developed by specific companies. These include Anthropic's Claude and OpenAI's GPT series. The cost per token is significantly higher than open alternatives. This high price applies to every single task the agent performs.
The specific open models that replaced expensive closed ones
Open models have meaningfully closed the gap between top models from closed labs. Leaders like Kimi K3 and GLM 5.3 handle many hard coding tasks. They do this at a fraction of the per-token price. Small models like GLM 5.3 Flash and DeepSeek V4.1 Flash handle everyday tasks well. These smaller models are capable enough for routine development work.
Together Link brings these specific open models into your existing tools. It supports a wide range of coding agents currently in use. The platform curates over 40 models chosen specifically for production use. This curation ensures quality while maintaining lower costs than closed options. Teams get access to frontier capabilities without paying premium prices.
Google launched EmbeddingGemma 2, a compact open-weight model that maps text, images, and audio into one vector space. It runs locally on phones with minimal RAM and enables instant on-device semantic search.
Google launched EmbeddingGemma 2, a compact open model that handles text, code, images, video, and audio. It uses a single shared vector space to enable unified search across all media types.
Which coding agents work with Together Link today
Together Link supports several major coding agents available today. It works with Claude Code and Claude Desktop versions. The tool also integrates with Codex in the ChatGPT app. Codex exists as a CLI application that benefits from this integration. OpenCode and Pi are other supported coding agent platforms. Your team keeps working normally with these specific agents.
Settings and logins stay exactly as they were before the switch. There is nothing new for users to learn about the system. Going back to native closed models takes only one command. This flexibility allows teams to test open models safely. They can revert to paid models if a task requires it.
How the Auto routing system picks models for tasks
The Auto routing system uses the Together AI Router to pick models automatically. It reads each session's first task to decide which model to use. Quick fixes go to fast, low-cost models automatically. Hard problems get assigned to frontier capability models instead. This decision happens once per session only.
Prompt caching keeps working because routing is efficient. The system routes between Opus 5.5 and GLM 5.3 if you have an Anthropic key. Without an Anthropic key, it routes between GLM 5.3 and GLM 5.3 Flash. This logic ensures the best model fits every specific task type.
Why this matters for cost, speed, and quality
This change matters because engineering spend drops by over 50%. Coding agents previously ran on expensive closed models constantly. Open models now provide similar quality at a lower price point. Speed remains competitive across the board of supported models. Billing runs on existing API keys without new contracts.
Quality does not suffer when switching to these open alternatives. Frontier models like Kimi K3 match performance levels of top closed labs. The per-token price is much lower for these open options. Small models handle everyday tasks efficiently and quickly. Teams save money while maintaining high productivity standards.
How it compares - what existed before, what this changes and what stays the same
Before Together Link, teams chose models manually. They often picked one model for all tasks. This approach ignores task difficulty levels. Some tasks need speed while others need deep reasoning. Closed models like Opus 5.5 handled everything but cost heavily. Open models offer better value but lacked easy integration. Teams had to manage multiple accounts and keys.
Together Link changes how teams select models automatically. It routes tasks to the best model instantly. The system analyzes the first task of each session. It then assigns a model based on complexity. Quick fixes get fast, cheap models immediately. Complex problems get powerful frontier models instead. This logic replaces manual selection entirely.
Settings and logins remain unchanged for users. Teams do not need new credentials or accounts. The workflow looks identical to before the switch. Users can revert to closed models easily. One command brings back native models. This flexibility reduces adoption friction significantly.
The infrastructure remains serverless based. Together AI provides the underlying compute power. Developers already trust this platform for inference. OpenRouter uses similar routing logic often. Together Link integrates directly with OpenRouter standards. This ensures compatibility with existing toolchains.
Questions this leaves open - what the source does not say and how a reader can check it
The article mentions specific models but lacks benchmark scores. Readers need to verify performance differences themselves. Independent benchmarks show actual token counts per minute. Teams should run their own tests on critical tasks. They must measure latency for real-world scenarios.
Cost savings depend heavily on usage volume. A small team might not see big reductions. Large teams with high token counts save more money. The 50% figure applies to specific workloads only. Engineers need to track their own spend carefully. They should monitor per-session costs over time.
Prompt caching behavior depends on session structure. Long sessions might benefit more from caching. Short bursts of activity waste cache opportunities. Teams should optimize their prompt strategies too. Repeated tasks in a session save extra money.
Security concerns exist regarding open model providers. Open models do not have the same security guarantees as closed ones. Some organizations require strict data privacy policies. Together AI must prove its security standards match competitors. Audits and compliance reports remain important for enterprises.
Integration depth varies across different agent platforms. Support lists include major coding agents today. Smaller or newer tools might lack full support yet. Teams should check compatibility lists regularly. Documentation updates reflect new integrations quickly.
Model availability changes based on demand. Frontier models can face capacity limits during peaks. Queues might form if usage spikes suddenly. Teams need contingency plans for model unavailability. Fallback mechanisms ensure work continues smoothly.
The article does not mention long-term cost trends. Prices fluctuate with market conditions constantly. Inflation and supply chain issues affect costs too. Engineers should review pricing pages monthly. They must adjust budgets accordingly over time.
Evaluation metrics like accuracy vary by task type. Coding tasks require different success measures than writing. Teams need tailored evaluation frameworks for each use case. Automated testing suites provide objective quality data.
Together AI's serverless platform scales dynamically. This means costs rise with demand automatically. Predicting exact costs requires historical usage data. Machine learning models forecast future spending accurately.
The routing logic prioritizes cost over speed sometimes. Speed might suffer if a cheaper model is slower. Teams need to balance cost and performance carefully. Some tasks prioritize latency above all else.
Open weights availability varies by region. Some models are not available globally yet. Local deployment options remain limited for now. Cloud providers host most open model instances.
Together Link supports specific coding agents listed. Other platforms might require custom integration work. Teams should verify their stack compatibility first. Missing support creates extra setup overhead.
The article lacks details on error handling rates. System failures can disrupt agent workflows unexpectedly. Downtime impacts productivity and project timelines. Reliability metrics matter for mission-critical operations.
Data retention policies for sessions are unclear. Teams need to know how long logs persist. Privacy regulations dictate data storage requirements. Compliance teams must review these policies closely.
The pricing model uses serverless pay-as-you-go. This means no upfront capital expenditure required. However, variable costs can surprise budget planners. Cash flow management becomes more complex.
Credit packs offer predictable billing structures. Fixed amounts prevent unexpected bill shocks. Teams might prefer credit packs for planning. They provide clearer financial visibility than usage-based pricing.
The article stops before discussing migration timelines. How long does it take to switch fully? Weeks or months depend on team size. Phased rollouts reduce risk during transitions.
Training new engineers on the tool takes time. Documentation quality affects learning curves significantly. Teams need clear guides for onboarding. Training sessions ensure everyone understands the system.
Feedback loops from users improve the product over time. User reports help developers fix bugs quickly. Community forums discuss issues and solutions openly.
Together AI's roadmap influences future feature additions. New models arrive based on development cycles. Engineers should stay updated on releases. Version compatibility matters for long-term stability.
The article does not mention enterprise support levels. Dedicated account managers handle large contracts. Small teams get general email support only. Response times vary by contract size and type.
SLAs (Service Level Agreements) define uptime guarantees. Downtime penalties apply if targets are missed. Teams should negotiate SLAs for critical workloads. Contracts protect against service disruptions.
Model drift occurs when performance degrades over time. Regular retraining keeps models accurate and relevant. Data pipelines feed fresh information into systems. Stale models produce poor results eventually.
The article lacks details on model versioning strategies. Older versions might remain available for legacy tasks. New versions introduce breaking changes sometimes. Teams need upgrade paths planned carefully.
The article mentions "frontier quality" but defines it vaguely. What exactly does frontier mean in this context? Specific benchmarks define the quality threshold. Teams need clear definitions for their standards.
Together Link's API documentation remains incomplete in some areas. Developers need full reference guides for integration. Missing endpoints cause implementation roadblocks. Documentation updates reflect new capabilities quickly.
The article does not discuss regional data residency laws. Some regions require data to stay local. Cross-border transfers trigger compliance review processes. Legal teams must evaluate these constraints.
Model licensing terms vary by provider and region. Open models often have different license types. Commercial use requires careful license review. Teams must check usage restrictions thoroughly.
The article mentions "over 50%" savings but not minimums. Some workloads might save less than half. Usage patterns determine the actual discount received. Baseline costs matter for calculating savings accurately.
Together AI's infrastructure uses specific hardware types. GPU models dominate current inference workloads. Chip availability affects pricing and performance. Hardware shortages can impact delivery times.
The article lacks details on model fine-tuning capabilities. Custom training extends model utility significantly. Fine-tuned versions offer specialized domain knowledge. Training costs add to total expenses.
Model selection involves trade-offs between speed and accuracy. Faster models often sacrifice some precision. Teams must choose based on task needs. Context windows limit input data sizes.
The article does not mention context window limits explicitly. Long documents exceed standard token limits. Summarization techniques handle large inputs better. Chunking strategies manage memory usage efficiently.
Together Link supports multiple model families today. DeepSeek, GLM, Kimi lead the current offerings. Other labs contribute models to the pool. Ecosystem diversity ensures broad coverage.
The article mentions "coding agents" but not specific tasks. Code generation, debugging, and refactoring all apply. Task categorization drives routing decisions effectively. Specificity improves model selection accuracy.
Model performance varies across programming languages. Some models excel at Python or Java. Others handle JavaScript better than others. Language-specific strengths guide model choices.
The article lacks details on API rate limits. Throttling occurs when requests exceed quotas. Quotas reset based on time windows. Teams need to plan request volumes carefully.
Together AI's billing system integrates with existing tools. Invoicing platforms connect for payment processing. Financial teams manage invoices and receipts easily. Automated reconciliation reduces manual work.
The article does not mention model explainability features. Black box models lack transparent reasoning paths. Explainable AI helps debug failures effectively. Debugging tools reveal model decision logic.
Model hallucinations remain a risk in open systems. False outputs occur without human verification. Fact-checking processes validate generated content reliably. Human-in-the-loop approaches mitigate risks.
The article mentions "serverless" but not cold start times. Latency spikes happen during initial loads. Warm-up periods stabilize performance quickly. Caching mitigates cold start effects.
Together Link's monitoring dashboard provides real-time metrics. Engineers track usage and costs live. Alerts notify teams of anomalies immediately. Proactive management prevents costly surprises.
The article does not discuss model selection algorithms in depth. How exactly does the router decide? Heuristics guide the automatic selection process. Machine learning optimizes routing decisions over time.
Model compatibility with existing workflows varies. Some pipelines require minor adjustments for open models. Legacy systems might need updates for support. Compatibility testing ensures smooth integration.
The article mentions "Auto" mode but not manual override options. Teams retain control to force specific models. Manual selection overrides automatic routing choices. Flexibility allows precise task management.
Together AI's community contributes model improvements constantly. Open source projects drive innovation forward. Community feedback shapes development priorities. Collaborative efforts accelerate progress globally.
The article lacks details on model security patches. Vulnerabilities require timely updates and fixes. Security teams monitor for threats regularly. Patch schedules protect against exploits.
Model performance degrades with increased token counts. Longer contexts consume more memory resources. Optimization techniques reduce memory footprint significantly. Efficient processing handles large inputs better.
The article does not mention model versioning history. Past versions remain accessible for reference tasks. Version tracking helps maintain consistency. Historical data aids debugging efforts.
Together Link's support team offers tiered assistance levels. Enterprise clients get priority response times. Standard users wait longer for replies. Support quality correlates with contract size.
The article mentions "frontier models" but not specific capabilities. What makes them frontier compared to others? Superior reasoning or coding skills define the group. Specific advantages distinguish top performers.
Model training data sources vary by provider. Training sets influence model knowledge boundaries. Data freshness impacts information accuracy levels. Outdated data produces irrelevant responses.
The article lacks details on model evaluation frameworks. How do teams measure success objectively? Automated tests provide quantitative metrics. Human evals offer qualitative insights too.
What to do to set up Together Link in minutes
You can set up Together Link by running a single command. Use the curl command to install the tool on your machine. The script downloads the installer from link.together.ai automatically. Your agent opens exactly as before after installation completes. You must create an account for an API key first.
Use "Auto" to let the router pick the best model for you. See your savings after each session ends quickly. A per-session tracker shows what you spent versus Opus 5.5 costs. Billing runs against your existing Together API key directly. Choose serverless pay-as-you-go or credit packs for billing.