The brief
- Sohu ASIC: 500,000 tokens/sec on Llama 70B; 20x throughput vs. 8x H100 GPUs.
- Built on TSMC N4P; proprietary low-voltage (LVI) and cluster-scale memory (CSM) techniques.
- August 2026: Shipped first production racks to Jane Street; $21B valuation after $700M Series D.
Technical approach
Single-architecture design
Optimized for transformer inference only
Built for decoding tokens only. Instruction sets and memory paths designed purely for this task.
Low-voltage inference (LVI)
Math at half the voltage of competing chips
Runs floating-point ops at <50% voltage vs. other AI chips. Cuts power and heat.
Cluster-scale memory (CSM)
Shared memory with SRAM latency
Multiple chips share memory via proprietary interconnects with SRAM-level latency.
