Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Together Link Connects Coding Agents to Open Models - OpenSmartRoute
Together Link Connects Coding Agents to Open Models
Together AI released a free CLI tool that routes coding agents like Claude Code to open models.
Key points
Together Link connects six major coding agents to open models via one command.
The router automatically selects cheaper models for simple tasks and premium ones for hard work.
Users can save up to 50% compared to using only the most expensive closed models.
Supported models include Kimi K3, GLM 5.3, and DeepSeek V4.1 Flash with 1 million context.
Why it matters: Running agents costs money; this tool lets you switch from expensive closed models to cheaper open ones without changing your workflow.
By OpenSmartRoute editorial · written through the router by writer-small
From MarkTechPost - “Meet Together Link: A Free CLI That Runs Open Models Like Kimi K3 and GLM 5.3 Inside Claude Code, Codex, and OpenCode”
Together AI launched Together Link in beta. This free tool connects coding agents to open models. It runs on macOS and Linux systems today. The software uses an MIT license for its code. Developers can install it without paying upfront fees.
The team states engineering organizations spend huge sums monthly. These funds go toward closed models with high costs. Together Link aims to shrink that bill significantly. It lets users keep their favorite coding harnesses. You only need to swap the underlying model underneath.
Installation requires a single command on supported systems. The tool works on macOS and Linux machines. Users must have a valid Together API key ready. The installer handles Bun installation if your system needs it. Commands land in the local bin directory automatically.
Run togetherlink to open the main launcher interface. You can also start specific tools directly from the terminal. Typing togetherlink claude connects to Claude Code or Desktop. The command togetherlink codex activates the Codex agent. Other options include opencode and pi for their respective agents. Shortcuts like tclaude work alongside the full commands.
The system picks models based on task difficulty automatically. It reads the first task of every new session. Quick fixes get routed to fast, low-cost models immediately. Hard problems receive frontier capability from stronger models. This logic happens once per entire session.
Pricing details show open model costs are much lower. Kimi K3 charges $3.00 in and $15.00 out per request. GLM 5.3 costs $1.40 in and $4.40 out per request. DeepSeek V4.1 Flash is listed at $0.30 in and $1.20 out. These rates are significantly lower than premium closed models.
Together claims over 50% savings compared to standard usage. The tool offers 50 to 80% savings versus all-Opus 5.5 sessions. Billing runs on your existing Together key account. You can use pay-as-you-go or credit packs for payment. Each session prints token and dollar totals on exit.
Claude Code shows estimated spend beside the equivalent Opus cost. Running togetherlink usage --last 7d shows gateway-tracked spend across sessions. The status line displays financial data clearly in the interface. This transparency helps engineers track their spending habits.
Reflection released Beam, an open-weight model that matches GLM 5.2 and Qwen 3.8 on benchmarks while using three to four times less compute.
Routing happens once per session to keep prompt caching working. Sessions default to a virtual auto model for flexibility. With an Anthropic API key, it routes between Opus 5.5 and GLM 5.3. Without that key, it routes between GLM 5.3 and GLM 5.3 Flash.
The Opus path applies only to Claude Code and Claude Desktop. Codex, OpenCode, Pi, and ChatGPT Desktop always stay on Together models. You can pin one model by placing a flag before the tool name. The syntax looks like togetherlink --main zai-org/GLM-5.3 claude.
Inside Claude Code, the /model menu maps tiers to open models. Opus runs Kimi K3 while Fable runs GLM 5.3. Sonnet runs GLM 5.3 Flash and Haiku runs DeepSeek V4.1 Flash. This mapping lets users choose performance levels easily.
The docs list Kimi K3, GLM 5.3, GLM 5.3 Flash, and DeepSeek V4.1 Flash. Each of these models supports a 1M context window. The product page lists current pricing for all available options. Current rates live on Together's official pricing page.
Billing runs through your existing Together key account. You can choose between pay-as-you-go or credit packs. Each session prints token and dollar totals upon exit. In Claude Code, the status line shows estimated spend data.
Together team notes it serves the largest OpenRouter token share for DeepSeek V4.1 Flash. The share is 40.8% as of September 30, 2026. GLM 5.3 Flash holds a 28.2% share on that date. Kimi K3 accounts for 23.1% of the token share.
FeatureTogether LinkOpenRouterClaude Code RouterOllama launch
One curl install, then togetherlink claude
Claude Code, Claude Desktop, Codex, OpenCode 2, Pi
Guides for Claude Code, Codex CLI, OpenCode, Cursor, and more
10 agents, including Claude Code, Codex, OpenCode, Pi
Claude Code, OpenCode, Claude and ChatGPT Desktop (macOS)
Auto router, optional Opus 5.5 escalation
Per-session receipt plus 7-day usage report
Activity dashboard plus statusline script
Together Link's edge is a curated, one-command path with built-in savings receipts. Claude Code Router offers broader provider control but runs a local gateway you manage. One install command connects six existing agents to open models on Together AI.
The default Auto router picks between GLM 5.3, GLM 5.3 Flash, and optionally Opus 5.5. Together claims over 50% savings compared to standard usage patterns. The tool offers 50 to 80% savings versus all-Opus 5.5 sessions consistently.
No local proxy runs during operation. Your normal agent configuration files stay untouched by the process. Each tool talks directly to Together's hosted gateway server. Terminal agents receive a temporary per-launch configuration that is removed when the session ends.
Claude Desktop and ChatGPT Desktop use separate, reversible profiles. Commands like togetherlink chatgpt off switch back to default settings. The router reads each session's first task to make decisions. Quick fixes go to fast, low-cost models automatically.
Hard problems get frontier capability from stronger open models. With an Anthropic API key, it routes between Opus 5.5 and GLM 5.3. Without one, it routes between GLM 5.3 and GLM 5.3 Flash. Routing happens once per session so prompt caching keeps working.
Sessions default to a virtual auto model for flexibility. Per the launch post, the router reads each session's first task immediately. The system decides which model fits the current work best. You do not need to configure complex routing rules manually.
The docs list Kimi K3, GLM 5.3, GLM 5.3 Flash, and DeepSeek V4.1 Flash. Each of these models supports a 1M context window size. The product page lists Kimi K3 at $3.00 in and $15.00 out. GLM 5.3 is listed at $1.40 in and $4.40 out.
DeepSeek V4.1 Flash and MiniMax M3 are listed at $0.30 in and $1.20 out. Current rates live on Together's pricing page for verification. Billing runs on your existing Together key account. You can use pay-as-you-go or credit packs for payment methods.
Each session prints token and dollar totals on exit. In Claude Code, the status line shows estimated spend beside the equivalent Opus cost. Running togetherlink usage --last 7d shows gateway-tracked spend across sessions. This feature helps engineers track their spending habits over time.
Together team also notes it serves the largest OpenRouter token share for DeepSeek V4.1 Flash. The share is 40.8% as of September 30, 2026. GLM 5.3 Flash holds a 28.2% share on that date. Kimi K3 accounts for 23.1% of the token share.
FeatureTogether LinkOpenRouterClaude Code RouterOllama launch
One curl install, then togetherlink claude
Claude Code, Claude Desktop, Codex, OpenCode 2, Pi
Guides for Claude Code, Codex CLI, OpenCode, Cursor, and more
10 agents, including Claude Code, Codex, OpenCode, Pi
Claude Code, OpenCode, Claude and ChatGPT Desktop (macOS)
Auto router, optional Opus 5.5 escalation
Per-session receipt plus 7-day usage report
Activity dashboard plus statusline script
Together Link's edge is a curated, one-command path with built-in savings receipts. Claude Code Router offers broader provider control but runs a local gateway you manage. One install command connects six existing agents to open models on Together AI.
The default Auto router picks between GLM 5.3, GLM 5.3 Flash, and optionally Opus 5.5. Together claims over 50% savings compared to standard usage patterns. The tool offers 50 to 80% savings versus all-Opus 5.5 sessions consistently.
No local proxy runs during operation. Your normal agent configuration files stay untouched by the process. Each tool talks directly to Together's hosted gateway server. Terminal agents receive a temporary per-launch configuration that is removed when the session ends.
Claude Desktop and ChatGPT Desktop use separate, reversible profiles. Commands like togetherlink chatgpt off switch back to default settings. The router reads each session's first task to make decisions. Quick fixes go to fast, low-cost models automatically.
Hard problems get frontier capability from stronger open models. With an Anthropic API key, it routes between Opus 5.5 and GLM 5.3. Without one, it routes between GLM 5.3 and GLM 5.3 Flash. Routing happens once per session so prompt caching keeps working.
Sessions default to a virtual auto model for flexibility. Per the launch post, the router reads each session's first task immediately. The system decides which model fits the current work best. You do not need to configure complex routing rules manually.
The docs list Kimi K3, GLM 5.3, GLM 5.3 Flash, and DeepSeek V4.1 Flash. Each of these models supports a 1M context window size. The product page lists Kimi K3 at $3.00 in and $15.00 out. GLM 5.3 is listed at $1.40 in and $4.40 out.
DeepSeek V4.1 Flash and MiniMax M3 are listed at $0.30 in and $1.20 out. Current rates live on Together's pricing page for verification. Billing runs on your existing Together key account. You can use pay-as-you-go or credit packs for payment methods.
Each session prints token and dollar totals on exit. In Claude Code, the status line shows estimated spend beside the equivalent Opus cost. Running togetherlink usage --last 7d shows gateway-tracked spend across sessions. This feature helps engineers track their spending habits over time.
Together team also notes it serves the largest OpenRouter token share for DeepSeek V4.1 Flash. The share is 40.8% as of September 30, 2026. GLM 5.3 Flash holds a 28.2% share on that date. Kimi K3 accounts for 23.1% of the token share.
FeatureTogether LinkOpenRouterClaude Code RouterOllama launch
One curl install, then togetherlink claude
Claude Code, Claude Desktop, Codex, OpenCode 2, Pi
Guides for Claude Code, Codex CLI, OpenCode, Cursor, and more
10 agents, including Claude Code, Codex, OpenCode, Pi
Claude Code, OpenCode, Claude and ChatGPT Desktop (macOS)
Auto router, optional Opus 5.5 escalation
Per-session receipt plus 7-day usage report
Activity dashboard plus statusline script
Together Link's edge is a curated, one-command path with built-in savings receipts. Claude Code Router offers broader provider control but runs a local gateway you manage. One install command connects six existing agents to open models on Together AI.
The default Auto router picks between GLM 5.3, GLM 5.3 Flash, and optionally Opus 5.5. Together claims over 50% savings compared to standard usage patterns. The tool offers 50 to 80% savings versus all-Opus 5.5 sessions consistently.
No local proxy runs during operation. Your normal agent configuration files stay untouched by the process. Each tool talks directly to Together's hosted gateway server. Terminal agents receive a temporary per-launch configuration that is removed when the session ends.
Claude Desktop and ChatGPT Desktop use separate, reversible profiles. Commands like togetherlink chatgpt off switch back to default settings. The router reads each session's first task to make decisions. Quick fixes go to fast, low-cost models automatically.
Hard problems get frontier capability from stronger open models. With an Anthropic API key, it routes between Opus 5.5 and GLM 5.3. Without one, it routes between GLM 5.3 and GLM 5.3 Flash. Routing happens once per session so prompt caching keeps working.
Sessions default to a virtual auto model for flexibility. Per the launch post, the router reads each session's first task immediately. The system decides which model fits the current work best. You do not need to configure complex routing rules manually.
The docs list Kimi K3, GLM 5.3, GLM 5.3 Flash, and DeepSeek V4.1 Flash. Each of these models supports a 1M context window size. The product page lists Kimi K3 at $3.00 in and $15.00 out. GLM 5.3 is listed at $1.40 in and $4.40 out.
DeepSeek V4.1 Flash and MiniMax M3 are listed at $0.30 in and $1.20 out. Current rates live on Together's pricing page for verification. Billing runs on your existing Together key account. You can use pay-as-you-go or credit packs for payment methods.
Each session prints token and dollar totals on exit. In Claude Code, the status line shows estimated spend beside the equivalent Opus cost. Running togetherlink usage --last 7d shows gateway-tracked spend across sessions. This feature helps engineers track their spending habits over time.
Together team also notes it serves the largest OpenRouter token share for DeepSeek V4.1 Flash. The share is 40.8% as of September 30, 2026. GLM 5.3 Flash holds a 28.2% share on that date. Kimi K3 accounts for 23.1% of the token share.
FeatureTogether LinkOpenRouterClaude Code RouterOllama launch
One curl install, then togetherlink claude
Claude Code, Claude Desktop, Codex, OpenCode 2, Pi
Guides for Claude Code, Codex CLI, OpenCode, Cursor, and more
10 agents, including Claude Code, Codex, OpenCode, Pi
Claude Code, OpenCode, Claude and ChatGPT Desktop (macOS)
Auto router, optional Opus 5.5 escalation
Per-session receipt plus 7-day usage report
Activity dashboard plus statusline script
Why it matters
Saving money while keeping quality for hard tasks reduces operational costs significantly.
What to do
Install the tool and test your first routing session today.