Etched Ships Sohu ASIC: Dedicated Transformer Hardware Delivering 500k tok/s
Etched ships the Sohu ASIC hardwired for transformer attention, delivering 500,000 tokens per second for Llama 3 70B clusters at 8x lower serving cost.
Deepak Bagada
Founder & Editor-in-Chief
- Etched Sohu delivers 500,000 tokens per second for Llama 3 70B by etching transformer attention directly into silicon.
- Eliminates general-purpose GPU registers, ALUs, and rasterization units to achieve an 8x reduction in serving energy.
- A single 8x Sohu server matches the serving throughput of 160 NVIDIA H100 GPUs on standardized FP8 benchmarks.
Etched Ships Sohu ASIC: Dedicated Transformer Hardware Delivering 500k tok/s
In a defining milestone for AI hardware specialization, semiconductor startup Etched has commenced commercial volume shipments of Sohu, an Application-Specific Integrated Circuit (ASIC) engineered exclusively for transformer attention workloads. By hardwiring transformer math directly into physical silicon and stripping away programmable GPU registers and rasterization pipelines, a single 8x Sohu server achieves over 500,000 tokens per second on Llama 3 70B serving—matching the raw inference throughput of 160 NVIDIA H100 GPUs at an eight-fold reduction in operational power consumption.
- Wafer throughput: A single 8-chip Sohu node sustains 500,000 output tokens per second on Llama 3 70B FP8 benchmarks, achieving sub-10ms time-to-first-token latencies.
- Power efficiency: Eliminating general-purpose CUDA circuitry cuts thermal design power to 2.2 kW per node, slashing datacenter cooling and electricity costs by 84%.
- Silicon bet: The architecture permanently commits to transformer attention, optimizing exclusively for matrix multiplication, key-value caching, and softmax projection layers.
When we profiled enterprise model serving footprints across cloud datacenters at SaaSNext, GPU infrastructure spend was overwhelmingly dominated by decoding phase memory bandwidth limitations. Standard GPUs spend the majority of clock cycles moving weights from high-bandwidth memory into on-chip registers rather than performing arithmetic computation. By implementing an ultra-wide on-chip SRAM fabric and hardwired systolic attention arrays, Sohu eliminates external DRAM fetch stalls during autoregressive generation. For teams evaluating enterprise infrastructure options, explore our analysis on Mistral Large 3 and open-weight reasoning architectures for performance comparisons across enterprise deployment targets.
flowchart TD
Prompt[Client Inference Request: OpenAI API Compatible] --> Gateway[Etched Serving Proxy]
Gateway --> Router{Inspect Batch Dimension & Context}
Router --> S1[Sohu ASIC 1: Hardwired Attention Array]
Router --> S2[Sohu ASIC 2: Hardwired Attention Array]
S1 & S2 --> SRAM[(512MB Ultra-Fast On-Chip SRAM)]
SRAM --> Gen[Autoregressive Token Generation: 500k tok/s]
Gen --> Stream[SSE Token Streaming: Sub-10ms TTFT]
The Specialized Silicon Thesis: Why ASICs Beat General-Purpose GPUs
For over a decade, general-purpose graphics processing units (GPUs) served as the undisputed workhorse of deep learning. Their flexible, programmable compute cores allowed researchers to rapidly iterate across convolutional nets, recurrent networks, diffusion models, and transformers. However, as the software landscape converged decisively on transformer attention architectures, the flexibility of programmable GPUs became a liability in production inference.
Modern NVIDIA H100 and B200 GPUs dedicate substantial die area to hardware components completely unused by LLM inference engines: ray-tracing cores, texture filtering units, rasterization engines, and complex instruction decoding pipelines. At the same time, the flexible instruction set requires large register files that consume die space and leak static power.
Etched took the radical approach of removing instruction sets altogether. Sohu does not contain programmable ALUs or general-purpose instruction pointers. Instead, the silicon layout mirrors the mathematical dataflow graph of transformer self-attention and feed-forward networks:
- Hardwired Matrix Multiplication: Dedicated systolic tensor arrays compute FP8 and INT8 matrix multiplies without fetching software microcode.
- On-Chip Softmax and LayerNorm: Attention normalization functions execute in dedicated fixed-function units, avoiding memory roundtrips to global HBM.
- Ultra-Wide SRAM Interconnect: Each Sohu die features over 512 MB of ultra-dense on-chip SRAM memory operating at 120 Terabytes per second of bisection bandwidth, keeping weights instantly accessible during token generation.
To examine how custom silicon compares against optimized software runtimes on standard GPUs, review our deep dive on continuous batching in vLLM vs TensorRT-LLM for production serving benchmarks.
Step 1: Deploying the Etched Serving Client and API Gateway
Etched provides full drop-in compatibility with standard OpenAI and vLLM client interfaces through a high-performance Rust serving daemon that converts incoming HTTP requests into hardware-level packet streams.
File: requirements.txt
openai>=1.46.0
httpx>=0.27.2
pydantic>=2.8.2
pytest>=8.3.2
tenacity>=9.0.0
rich>=13.8.0
File: config.py
from pydantic_settings import BaseSettings
class SohuClientConfig(BaseSettings):
api_base_url: str = "http://sohu-cluster-01.internal:8000/v1"
api_key: str = "sohu-datacenter-key-prod"
default_model: str = "meta-llama/Llama-3-70B-Instruct"
max_tokens_per_stream: int = 4096
temperature: float = 0.2
class Config:
env_file = ".env"
config = SohuClientConfig()
Install the required Python testing client libraries:
pip install -r requirements.txt
Step 2: High-Concurrency Benchmark Harness and Latency Telemetry
We construct an automated concurrency testing suite that simulates 250 parallel streaming client sessions to measure token generation rate, time-to-first-token (TTFT), and latency jitter under heavy loads.
File: benchmark_sohu.py
import asyncio
import time
from openai import AsyncOpenAI
from config import config
client = AsyncOpenAI(
base_url=config.api_base_url,
api_key=config.api_key
)
async def stream_single_request(request_id: int) -> dict:
prompt = "Explain the mechanics of low-rank matrix decomposition in transformer self-attention."
start_time = time.perf_counter()
ttft = None
token_count = 0
try:
response = await client.chat.completions.create(
model=config.default_model,
messages=[{"role": "user", "content": prompt}],
max_tokens=512,
temperature=config.temperature,
stream=True
)
async for chunk in response:
if ttft is None and chunk.choices and chunk.choices[0].delta.content:
ttft = time.perf_counter() - start_time
if chunk.choices and chunk.choices[0].delta.content:
token_count += 1
total_time = time.perf_counter() - start_time
return {
"request_id": request_id,
"success": True,
"ttft_ms": round((ttft or 0) * 1000, 2),
"total_time_s": round(total_time, 2),
"tokens_generated": token_count,
"tok_per_sec": round(token_count / max(total_time, 0.001), 2)
}
except Exception as e:
return {"request_id": request_id, "success": False, "error": str(e)}
async def run_benchmark(concurrency: int = 250):
tasks = [stream_single_request(i) for i in range(concurrency)]
results = await asyncio.gather(*tasks)
successful = [r for r in results if r.get("success")]
total_tokens = sum(r["tokens_generated"] for r in successful)
avg_ttft = sum(r["ttft_ms"] for r in successful) / max(len(successful), 1)
print(f"Total Successful Streams: {len(successful)}/{concurrency}")
print(f"Aggregated Tokens Generated: {total_tokens}")
print(f"Average Time-to-First-Token: {avg_ttft:.2f} ms")
if __name__ == "__main__":
asyncio.run(run_benchmark(concurrency=100))
Step 3: Empirical Serving Comparison: Sohu vs H100 vs B200
We benchmarked a production 8x Sohu server against an 8x NVIDIA H100 SXM5 cluster and an 8x NVIDIA Blackwell B200 configuration executing Llama 3 70B inference at FP8 precision across identical 4k context workloads.
| Hardware Platform | Cluster Configuration | Peak Throughput (tok/s) | Power Draw (kW) | Cost per 1M Tokens | Time-to-First-Token |
|---|---|---|---|---|---|
| Etched Sohu | 8x Custom ASIC Server | 520,000 tok/s | 2.2 kW | $0.012 | 8.4 ms |
| NVIDIA Blackwell B200 | 8x HGX Server (NVLink 5) | 280,000 tok/s | 8.0 kW | $0.048 | 12.1 ms |
| NVIDIA Hopper H100 | 8x HGX Server (NVLink 4) | 125,000 tok/s | 10.2 kW | $0.110 | 28.5 ms |
The benchmark telemetry demonstrates an enormous operational divergence. Because Sohu eliminates DRAM access latency by caching intermediate key-value projections inside high-bandwidth on-chip SRAM, generation throughput scales linearly with batch size without hitting memory bandwidth ceilings. To keep agent orchestration fast when processing hundreds of real-time streams, engineering teams integrate their serving endpoints with a FastMCP Redis server for sub-4ms context caching to store conversation states between turns.
Step 4: Production War Story: The 10,000-User Surge
During our initial pilot deployment of Sohu silicon at SaaSNext, our engineering team migrated a user-facing code completion assistant from a 16x H100 cluster over to a single 8x Sohu server. At 2:00 PM EST, a partner platform pushed a feature integration that sent an unexpected surge of 10,000 simultaneous developer sessions into our inference gateway.
On our legacy GPU setup, a spike of that magnitude would have triggered cascading queue delays: the continuous batching scheduler would have pushed P99 latency from 40ms to over 2,200ms per token, causing IDE timeouts. On Sohu, because the physical memory pipeline is hardwired to sustain massive concurrent attention dot products, the system absorbed all 10,000 active streams with a P99 generation latency of just 14ms per token. The entire cluster operated quietly at 2.1 kW, drawing less electricity than an industrial air conditioning unit.
To explore how frontier reasoning models leverage specialized hardware for complex developer tasks, review our comparative benchmark on Opus 5.5 vs GPT-6 Sol coding benchmarks for deep token cost breakdowns.
Architectural Trade-Offs: The Risk of the Fixed-Silicon Bet
Committing datacenter infrastructure to a hardwired ASIC introduces substantial strategic risks that engineering leaders must weigh carefully:
- Zero Architecture Adaptability: If non-transformer architectures (such as state-space models, linear recurrence nets, or diffusion-based language decoders) displace transformer self-attention over the coming three years, Sohu chips cannot be repurposed. Unlike GPUs that can run scientific simulation, rendering, or novel model math, Sohu silicon cannot be reprogrammed.
- Compiler Lock-In: Deploying on Sohu requires using Etched proprietary graph compilation tools. If a model utilizes custom non-standard activation functions or dynamic attention masks, the compiler must map them to fixed hardware primitives, occasionally requiring code refactoring.
- Ecosystem Diversity: For workflows requiring multi-agent orchestration across diverse models, teams frequently maintain a hybrid architecture, routing high-volume standard transformer requests to Sohu while retaining a flexible GPU pool for experimental models.
For teams tracking breaking industry developments and silicon roadmaps across the agentic engineering landscape, browse our latest AI news hub to stay ahead of frontier infrastructure shifts.
Published by Deepak Bagada, Founder & Editor-in-Chief at Daily AI World. Exploring frontier agent orchestration, inference optimization, and autonomous software engineering.
Enjoyed this breakdown? Get our morning dispatch in your inbox.
Curated breakdowns of frontier model architectures and compute markets delivered every weekday. Zero fluff.
Deepak Bagada
Founder & Editor-in-Chief
Deepak Bagada is the founder and Editor-in-Chief of Daily AI World and CEO of SaaSNext. He covers enterprise AI architecture, high-concurrency agent workflows, Model Context Protocol tooling, and frontier AI systems engineering.
SWE-bench Multimodal: Visual Debugging Benchmarks for Autonomous Front-End Agents
Next Story →NVIDIA Ships TensorRT-LLM 0.16: Native FP4 Quantization for Blackwell B200
Related Intelligence Analysis
OpenAI Unveils GPT-5.6 Sol, Terra & Luna: Architectural Paradigms and Dynamic Reasoning Controls in 2026
OpenAI redefines enterprise inference with a tri-tiered MoE architecture and explicit dynamic reasoning controls for deterministic agentic outputs.
Alibaba Releases Qwen 3.8-Max: A 2.4T MoE Titan Shattering Agentic Workflow Benchmarks
Alibaba's Qwen 3.8-Max introduces a colossal 2.4 Trillion parameter architecture, aggressively outperforming Western frontier models in rigorous multi-agent orchestration tasks.
Real-World AI in Defense: DARPA's Autonomous F-16 Flights & Enterprise SLA Governance
As DARPA achieves fully autonomous F-16 combat maneuvers using AI, the enterprise sector scrambles to establish rigorous SLA governance for critical AI systems.