Etched’s $10.3 Billion Bet: A Transformer-Only Chip Takes On AI Inference Dominance

Etched closed a $300M Series C at $10.3B valuation led by Sequoia with a16z and SK Hynix. Its Sohu ASIC hardcodes transformers for superior inference efficiency. With $1B in contracts and racks shipping this summer, the startup challenges Nvidia's dominance in a specialized hardware race. The bet carries high stakes but strong backers.
Etched’s $10.3 Billion Bet: A Transformer-Only Chip Takes On AI Inference Dominance
Written by Sara Donnelly

Etched just closed a $300 million Series C. The post-money valuation hit $10.3 billion. Sequoia led the round. Andreessen Horowitz, Jane Street, Diffusion and SK Hynix joined in. The numbers tell one story. The bet behind them tells another.

Three Harvard dropouts founded the company in 2022. Gavin Uberti serves as CEO. Chris Zhu acts as CTO. Robert Wachen is president. They dropped out of school after Peter Thiel backed them. Now their firm employs more than 400 engineers. Many come from Nvidia, Google TPUs, Broadcom, SK Hynix and TSMC. The talent pool runs deep. So do the ambitions.

The Hardware Gamble That Ignores Flexibility

Etched builds one thing. An application-specific integrated circuit that runs transformers and nothing else. The chip, called Sohu, hardcodes the attention mechanisms at the heart of nearly every major large language model. No general-purpose cores. No support for arbitrary computations. Just raw speed on the workloads that matter most right now.

Early customer tests show state-of-the-art throughput, latency and power efficiency. The company claims its low-voltage inference architecture runs math blocks at under half the voltage of most AI chips. That yields multiple times the FLOPs density. Sparse mixture-of-experts models with trillions of parameters hit 80%+ peak FLOPs. No thermal throttling. The numbers sound aggressive. Real silicon from TSMC’s N4P process backs them up. A0 chips returned earlier this year.

But power density solves only part of the problem. Memory bandwidth and latency create bigger headaches at cluster scale. Etched’s cluster-scale memory approach builds a shared memory pool across the entire scale-up domain. A proprietary ultra-low-latency, high-bandwidth interconnect combines HBM and SRAM in a hybrid design. Latency drops. Capacity grows. The usual trade-offs between SRAM-only chips, 3D DRAM stacks or optical links simply disappear. Or so the company asserts.

Co-design sits at the center of everything. Chips. Racks. Software. Manufacturing methods. All optimized together for prefill and decode phases of inference. The rack-scale product now sits in validation with customers. First racks ship this summer. Production has kicked off to fulfill over $1 billion in signed contracts. Demand already outstrips supply. Evaluation units move quickly to deployment. Fast.

Uberti captured the mood in recent comments. “Now time to be aggressive. Time to ship.” The message lands with force. Etched emerged from stealth in 2025 with working hardware and that initial $1 billion pipeline. It has scaled facilities since. An 80,000-square-foot site in San Jose houses a data center, test house and NPI prototyping lab. A factory in Taiwan complements the U.S. footprint. Vertical integration aims at gigawatt-scale output. Bold. Expensive. Necessary if the company wants to matter.

Investors bought the vision long before today’s round. The company raised roughly $800 million across earlier unannounced financings. Peter Thiel participated. So did Andrej Karpathy, Geoffrey Hinton, Tri Dao, Fei-Fei Li and a long list of AI luminaries and operators. Total capital now exceeds $1 billion. The $10.3 billion valuation more than doubled from a reported $5 billion earlier in 2026, according to Data Center Dynamics.

That earlier jump came after a $500 million round led by Stripes with Thiel’s involvement. Reuters noted this week’s Sequoia-led round marks the highest valuation for any Series C the firm has led. Momentum feels palpable. Skeptics remain. Can a single-architecture chip survive when models evolve? Etched insists transformers dominate for the foreseeable future. Mixture-of-experts variants, long-context workloads and agentic systems all map neatly onto its strengths. The company even claims architecture-agnostic elements in practice. Models such as DeepSeek, Qwen and even non-transformer architectures like Mamba reportedly run. Details stay sparse.

Competition tells its own tale. Nvidia still owns the lion’s share of AI accelerators. Its H100 and Blackwell parts power most training and inference clusters today. Hyperscalers build custom silicon too. Google deploys TPUs. Amazon offers Inferentia and Trainium. Meta works on MTIA. Yet none specialize so narrowly. None burn the transformer itself into silicon. Etched’s pitch rests on that difference. Specialization beats generality when the workload stays fixed. The math favors fixed-function hardware. Area, power and clock speeds improve dramatically without overhead for unused features.

Analysts watch the power and cooling equations closely. Inference clusters consume enormous electricity. Anything that cuts voltage and heat while raising effective FLOPs density carries obvious appeal. Etched’s rack-scale focus adds another layer. Customers buy full systems, not individual cards. Integration headaches shrink. Deployment accelerates. The $1 billion in contracts suggests some large buyers already agree.

Production realities loom large. TSMC manufactures the chips. SK Hynix supplies memory expertise and now capital. The new San Jose facility and Taiwan factory represent real capital expenditure. Scaling to compete with Nvidia’s volumes will test the team. Uberti knows the stakes. His background includes work on AI compilers and TVM. Zhu brings HPC research chops. Wachen co-founded earlier ventures that reached high valuations. Mark Ross, former CTO of Cypress Semiconductor acquired for $9.4 billion, now steers hardware execution. Brian Loiler spent 22 years at Nvidia building HGX and DGX platforms. The bench looks strong.

Yet questions persist. What happens when the next architectural wave arrives? Diffusion models, state-space models or something entirely new could shift priorities. Etched designed its hardware assuming transformer dominance for years. The bet carries risk. So far the market rewards it. Hyperscalers and AI labs hunt every efficiency gain. Token throughput per dollar and per watt decides winners in inference. Early benchmarks, though not fully public, reportedly show significant leads over H100-class GPUs on targeted tasks.

The company plans more performance disclosures this summer. Roadmap details will follow. Customers run terabytes of production traffic through simulators already. Dozens of engineers co-design directly with supply chain partners. The pace feels frantic. And deliberate.

Etched positions itself at the center of a broader shift. AI inference grows faster than training. Cost sensitivity rises. Power constraints tighten. Specialized hardware starts to look less like a niche and more like the logical next step. Google, Amazon and others made similar moves inside their clouds. Etched brings the idea to the broader market. Rack-scale products that ship this year could accelerate adoption.

Whether the valuation holds depends on execution. Shipping real racks. Hitting the claimed efficiency numbers at volume. Expanding the customer base beyond early adopters. The $10.3 billion price tag implies enormous expectations. Nvidia’s market cap sits orders of magnitude higher. But inference represents a massive slice of future spend. Capturing even a fraction with superior economics could justify the numbers.

Investors clearly believe. From Thiel to Karpathy to Stanley Druckenmiller, the cap table reads like a who’s who of AI and finance. Strategic backing from SK Hynix adds manufacturing credibility. Jane Street’s participation hints at quantitative rigor behind the efficiency claims.

So Etched races forward. Silicon validates. Racks assemble. Contracts pile up. The transformer, first proposed by Google researchers in 2017, now gets etched into physical circuits at scale. The circle feels complete. The question is whether this narrow bet rewrites the economics of AI deployment. Or becomes another cautionary tale of over-specialization.

For now the momentum runs one way. Summer shipments will deliver the first real verdict. Demand already exceeds supply. The company pushes production aggressively. Time to ship, indeed.

Subscribe for Updates

EmergingTechUpdate Newsletter

The latest news and trends in emerging technologies.

By signing up for our newsletter you agree to receive content related to ientry.com / webpronews.com and our affiliate partners. For additional information refer to our terms of service.

Notice an error?

Help us improve our content by reporting any issues you find.

Get the WebProNews newsletter delivered to your inbox

Get the free daily newsletter read by decision makers

Subscribe
Advertise with Us

Ready to get started?

Get our media kit

Advertise with Us