Articles Tagged “AI Inference”
9 articles found

Fujitsu MONAKA Puts 144 Arm Cores on Sale in November
Fujitsu's 2nm MONAKA server CPU packs 144 Armv9 cores and claims twice the AI inference throughput of rival CPUs at half the power. On sale in November.

d-Matrix Raptor XPUs Join Nvidia's NVLink Fusion Racks
d-Matrix will build its Raptor inference XPUs around Nvidia NVLink Fusion and MGX racks, with 3 TB/s per-XPU bandwidth and availability in Q4 2027.

WebGPU Kernels Make Local AI in the Browser 2.57x Faster
Hugging Face published 207 WebGPU kernels as a JavaScript library, reporting a 2.57x geometric-mean speedup over ORT WebGPU on an Apple M4 GPU.

Meta MTIA 400 Puts 9.4TB/s HBM3e Behind FP4 Inference
Meta detailed MTIA 400 at Hot Chips 2026: eight HBM3e stacks, 9.4TB/s of bandwidth, hardware FP4, and scale-up domains reaching 72 accelerators.

Cerebras CS-4 Packs Three Wafer-Scale Chips in a Rack
The Cerebras CS-4 puts three WSE-3 Turbo wafers in one rack for 750 PFLOPS, and Cerebras clocks 4,400 tokens per second per user on GPT-OSS-120B.

Etched Raises $700M at a $21B AI Chip Valuation
Etched closed a $700M Series D led by Jane Street at a $21 billion valuation, doubling in just a month as inference chip orders pass $1 billion.

Fireworks AI Raises $1.5B for Faster Model Inference
Fireworks AI closed a $1.505 billion Series D at a $17.5 billion valuation, a 4.4x step-up in nine months, betting on production model inference.

Orange Pi's AI Station Packs a Huawei Ascend 310 With 96GB RAM and 176 TOPS Into a Single Board Computer
The Orange Pi AI Station pairs 10 dedicated AI cores with 16 CPU cores, up to 96GB LPDDR4X, and 256GB eMMC — delivering data-center-class AI inference on a 130mm board.

Inception Labs Launches Mercury 2 — The First Reasoning LLM Built on Diffusion Architecture
Mercury 2 processes tokens in parallel via iterative denoising, hitting 1,000 tokens per second while matching top reasoning models on benchmarks.
