A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation

## Episode Summary In this episode, we cover: - **A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation** (arXiv) - **Performance Verification of the AmpereOne CPU Core** (arXiv) - **Intel says Nova Lake, its next major CPU architecture, is coming to desktop PCs before servers - TechSpot** (google_arch) - **Bit-Brick K1: Raspberry Pi 5 alternative with different CPU architecture, M.2 and PCIe support launches - Notebookcheck** (google_arch) - **India unveils a homegrown dual-core 1GHz RISC-V processor, the DHRUV64 - The Register** (google_riscv) --- *Sponsored by Ada, Ago Consulting, and Zen Semiconductor*