2026年9月10日

NVIDIA Invests $20B in Groq to Fill AI Inference Gap

Facing competition from Google’s TPU, NVIDIA CEO Jensen Huang decisively invested $20 billion to par...

Facing competition from Google’s TPU, NVIDIA CEO Jensen Huang decisively invested $20 billion to partner with hot chip startup Groq, aiming to strengthen AI inference capabilities. Technology investor Gavin Baker highlighted that GPUs face latency bottlenecks in inference: prefill stages suit GPU parallelism, but decode stages are sequential, and GPU performance is limited by HBM memory, leaving FLOPs underutilized.

Groq’s LPU chips leverage on-chip SRAM, running 100x faster than GPUs, handling 300–500 tokens per second at full load. However, each LPU only has 230MB of memory; large models require hundreds of LPUs, taking much more data center space than GPUs and incurring high overall hardware costs.

This strategic move allows NVIDIA not only to upgrade technology but also to defend its moat: integrating Groq’s low-latency inference chips addresses emerging market needs. The rise of TPU revealed GPU limitations in inference, and Groq provides NVIDIA a critical boost to maintain AI leadership.

接著讀

Microsoft Launches Another Major Layoff: 9,000 Jobs Cut Amid AI-Driven Workforce Restructuring

Following a 7,000-person layoff in May, Microsoft has announced another round of job cuts, planning to lay off about 9,000 employees—roughly 4% of its workforce. With AI being widely adopted in code generation and organizational management, the company is accelerating its structural overhaul. This marks one of Microsoft's largest workforce adjustments in recent years, reflecting how AI is reshaping roles across the tech industry.

432 天前