
StepFun Step 5 Preview: 600B MoE Agent Model at $1 Input
StepFun's Step 5 Preview pairs 600B parameters with 27B active and a 1M-token context at $1 per million input tokens, with open weights due October 15.
A Big Sparse Model With a Small Active Footprint
Shanghai-based StepFun has released Step 5 Preview, a sparse mixture-of-experts model built for long-horizon agent work, and put a date on its open weights. The API went live in the days before September 20, when MarkTechPost published its detailed breakdown; Artificial Analysis lists the release as September 18. The headline is an unusual pairing: a very large model that only wakes a small slice of itself for each token, sold at a price that makes long agent runs affordable.
- Architecture: 600B total parameters with roughly 27B active per token, across 92 transformer layers in a narrow-deep configuration
- Context: 1M tokens, with text, image and video input and text output
- Pricing: $1.00 per million input tokens, $0.05 for cached input and $2.70 per million output tokens
- Open weights: scheduled for October 15, 2026, needing about 1.2TB of storage in BF16
Why Does a 27B-Active MoE Matter?
The economics of a mixture-of-experts model come down to one ratio. Step 5 activates about 4.5% of its parameters for any given token, so it carries the knowledge capacity of a 600B model while paying the compute bill of something much smaller at inference time. If you want the full mechanics, our sparse mixture-of-experts explainer walks through how routers pick experts.
The narrow-deep design is the interesting choice here. Ninety-two layers is deep for a model of this class, and StepFun pairs it with MTP-3 speculative decoding, FP8 expert weights and KV-cache offload, all of which are serving optimizations aimed at keeping long, multi-step agent sessions fast and cheap. The 5-cent cached-input price is the tell: this model is priced for agents that re-read the same large context over and over.
How Does Step 5 Compare?
On independent measurement, Artificial Analysis gives Step 5 Preview a score of 44 on its Intelligence Index, placing it 24th, and notes that it is a verbose reasoning model that used a large number of tokens during evaluation. That verbosity is worth factoring into cost estimates, since output tokens are where the $2.70 rate applies.
StepFun's own numbers, run at its high effort setting, show 66.4 on FrontierFinance and 83.3 on the DRACO research benchmark. The company's comparison tables place those scores a few points below Claude Opus 5 and above GPT-6 Astra on the same tests, but those are vendor-reported figures and should be read as such until independent evaluations catch up.
What Can Developers Do With It Today?
Quite a lot. The API supports low, medium and high reasoning effort, tool calling, JSON mode and JSON Schema, streaming, and prompt caching. StepFun's documentation allows up to 60 images per request, and output runs up to 64K tokens, which suits report-style agent tasks.
Can You Run Step 5 Locally?
Not on a desktop. At roughly 1.2TB in BF16, self-hosting calls for multi-GPU server hardware, and even aggressive quantization will leave it well beyond a single consumer machine. Our LLM quantization guide explains what shrinking a model this size would involve. The real value of the October 15 release is for research labs and companies that want to fine-tune or host a capable agent model on their own infrastructure.
A firm open-weights date on a model of this scale is good news for the wider ecosystem. Follow the rest of this week's releases in our AI coverage.
Sources: Artificial Analysis — September 18, 2026; MarkTechPost — September 20, 2026; StepFun Platform Docs — September 2026.
More AI Stories

Strands Decider 2B: What Amazon's Free Decision Model Does
Strands Decider 2B is a free 2B-parameter decision model that picks options in about 115ms on a single GPU. Here is how it works and where it fits.

Meta Muse Gadgets SDK: Build Your Own AI Agent Hardware
Meta's open-source Muse Gadgets SDKs bring its Muse agent to Raspberry Pi and ESP32 builds, and 5,000 free Home Link dongles are going to subscribers.

AstaBrief 8B: Ai2's Open Model for Cited Science Reports
AstaBrief 8B is Ai2's Apache 2.0 open model that writes cited research reports in 51 seconds, 3.5x faster than its Claude pipeline. Here is how it works.
