Skip to content

Hardware

722 articles found

NVIDIA Launches Nemotron 3 Ultra: A 550B-Parameter AI Model Promising Frontier Reasoning at 30% Lower Cost

NVIDIA Launches Nemotron 3 Ultra: A 550B-Parameter AI Model Promising Frontier Reasoning at 30% Lower Cost

Jun 05, 2026
NVIDIA Technical Blog

NVIDIA launches Nemotron 3 Ultra, a massive 550B-parameter AI model delivering frontier reasoning capabilities at 30% lower cost and 5x higher throughput than comparable open models, powered by cutting-edge hybrid architecture and a novel multi-teacher training method, with fully open weights released for enterprise adoption.

New 'Wall Attention' Variant Delivers Per-Channel Forgetting Rates and Efficient Autoregressive Decoding With Full GQA Support

New 'Wall Attention' Variant Delivers Per-Channel Forgetting Rates and Efficient Autoregressive Decoding With Full GQA Support

Jun 03, 2026
GitHub

A new attention variant called Wall Attention launches with per-channel, per-timestep multiplicative decay for independent forgetting rates, backed by optimized Triton kernels enabling efficient autoregressive decoding, full GQA support, attention sinks, sliding windows, and sequence packing with verified numerical accuracy.

Previous
Page 17 of 73
Next
Showing 161 - 170 of 722 articles