Skip to content

PricingSaaS joins Willingness to PayRead the announcement →

Back
PRICING BRIEFINGDaily · Oct 2, 2026
TTogether
Together
PACKAGING CHANGE

Together adds Tev1 4B Experimental at $0.04 input with free output tokens

Together added Tev1 4B Experimental to its serverless inference pricing table at $0.04 per 1M input tokens, with output tokens priced at zero.

Together’s pricing page · Change Detected

WHY IT MATTERS

For developers evaluating Tev1 4B, the free-output pricing is unlikely to be permanent given Together's track record of raising rates on models once adopted — budget accordingly rather than assuming this rate holds. For the category, it's another data point in the broader move toward promotional, usage-based AI pricing that changes weekly rather than quarterly. Watch whether Tev1 4B gets its first price increase within the next month, following the same arc as Qwen3.7-Max.

BY THE NUMBERS

  • 38 real pricing changes since tracking began Dec 29, 2025 — about 1/week
  • 20 changes in the last 90 days alone
  • Previous change: Sep 25, 2026 — NVIDIA HGX B300 pricing reveal (7 days ago)
  • Qwen3.7-Max raised 3 times in September (up to 60% in one hike)

THEIR RATIONALE

Together ships pricing changes roughly once a week — 38 real changes in just over nine months of tracking, with 20 in the last 90 days alone — and adding a new experimental model at an aggressive promotional rate fits its pattern of using new-model launches to draw developers in before adjusting established models' prices. The 'Experimental' tag and free-output pricing suggest an introductory rate rather than a permanent one, consistent with how Together has previously launched models cheap and then repriced them within weeks, as it did with Qwen3.7-Max across three hikes in September alone.

TRACK RECORD38 changes since Dec 29, 202538 in the last 12 monthsa change every ~7 daysprevious change 7 days ago

THE SIGNAL

Together's tracked history skews heavily toward price increases and feature additions in near-equal measure — it adds models cheap, then raises rates once they prove popular, as it did with Qwen3.7-Max across three hikes in September. A new model priced at $0.04 input with free output sits at the cheap, loss-leading end of that cycle, not a stable rate-card entry. Zero-cost output tokens are rare enough in Together's own history that this reads as a deliberate adoption incentive for whatever workload Tev1 4B targets, rather than a reflection of Together's actual serving cost.

CREDITS & USAGE BILLING109 companies in the last 6 months11.8% of 1,151 tracked changes10× on Together's pageKoala most often (3)

IN CONTEXT

Together has made 38 real pricing changes since tracking began in December 2025 — a weekly cadence that makes this new-model addition unremarkable on its own, except for the pattern it fits. Together has repeatedly launched models cheap and then repriced them once adopted: Qwen3.7-Max jumped 60% on Sep 11, then another 25% ten days later, before Qwen3.8 Flash and Qwen3.7-Max were both cut 40% on Sep 23. Tev1 4B's $0.04-input, free-output launch rate looks like the same playbook's opening move. Credits-and-usage pricing is the single most common pattern in the index right now — 136 changes across 109 companies in six months, including Synthesia's shift to credit-based billing this week — so Together is moving with, not against, the broader market.

WHAT PEOPLE ARE SAYING

@aleexxx_aitwitter

A classifier that cost $17 to train and serves at $0.042/M input. That is the real story, not the model itself. Open weights are now cheap enough to iterate like software. Benchmarking this against a full-size classifier this week.

View original
@sandeepjtwitter

wow, $17 to train and $0/M output makes judge cost stop being a design constraint.... nice. The number I'd want before swapping: tev1-4B agreement with human labels vs a frontier judge on the same rubric. Cheap judging only helps if the agreement gap is small.

View original
@timniversetwitter

tev1-4B, together's jev-like classifier on qwen3.5 4B: ~$17 to train, ~25 min, 37,840 examples. $0.042/1M input, output free. the catch is buried in the tutorial: call the public endpoint directly and you have to set temperature=0, max_tokens=8 and thinking off yourself. it doesn't inject them

View original