2026年9月10日

NVIDIA’s AI Inference Strategy: Feynman + LPU, Blocking the ASIC Path?

NVIDIA is aiming to dominate the AI inference stack with its next-gen Feynman GPU, integrating Groq’...

NVIDIA is aiming to dominate the AI inference stack with its next-gen Feynman GPU, integrating Groq’s LPU units. According to GPU expert AGF, the LPU could be stacked on the Feynman chip using TSMC’s SoIC hybrid bonding, similar to AMD’s X3D CPU stacking 3D V-Cache.

The main advantage: LPU comes with large on-chip SRAM, enabling low-latency inference and high model floating-point utilization (MFU), while HBM remains for capacity storage. This provides NVIDIA a decisive edge in decoding and low-latency workloads, effectively limiting opportunities for ASIC competitors.

To avoid antitrust issues, NVIDIA did not acquire Groq outright. Instead, it signed a non-exclusive IP license, securing the rights to use LPU technology. This “reverse acquisition” mirrors Microsoft’s earlier strategy of absorbing startups’ talent and IP while staying outside regulatory scrutiny.

In short, the Feynman + LPU combo could boost NVIDIA’s AI inference capabilities to a new level, while making it much harder for ASIC rivals to compete.

接著讀