ELSEIF
Your brief EB
302 stories from 172 feeds 984 clusters Refreshed 4 minutes ago next pull 23:56

TOPIC

Performance

Profiling, benchmarks, and the hunt for the bottleneck that actually matters. Work that reports a baseline, a method, and a number you could reproduce yourself.

1TODAY
8FEEDS
4mMEDIAN
FEEDS Hacker News 55 Phoronix 40 Lobsters 24 Tomshardware 19 Techmeme 14 InfoQ 8 The New Stack 6 VentureBeat 4

PERFORMANCE

Everything in Performance.

02 367 new

Performance Techmeme

Apple unveils new Mac Studio with M5 Max and M5 Ultra

Why it matters — The update shifts the Mac Studio’s focus toward AI acceleration and memory capacity, making it a stronger option for local AI development and high-performance computing. However, the gains may not justify upgrades for users without AI-specific needs. Availability and pricing could limit adoption for smaller teams.

5 feeds
60 min
05 303 new

Performance malisper.me

pgrust JIT compiler compiles SQL queries in around 5μs using copy-and-patch

Why it matters — Traditional database JIT compilers rely on LLVM or C/C++ code generation, both of which have high compile times that limit when compilation is worthwhile. At 5μs, the compilation cost is low enough to apply JIT optimization to every query, including single-execution ones. The author also notes that AI assistance made directly targeting assembly far more approachable than historically expected.

3 feeds
15 min
06 299 new

Performance matklad

TigerStyle's static allocation and constant work patterns avoid pool use-after-free by pre-allocating all objects at startup

Why it matters — If you build systems with object pools, the type system does not track which generation of object occupies a slot, so a stale pointer can silently read bytes belonging to a different object. Pre-allocating a fixed maximum at startup and rejecting surplus requests trades a small loss of flexibility for a guarantee that overload degrades gracefully instead of triggering an OOM kill that loses every in-flight request.

3 feeds
7 min
07 270 new

Performance vectorware.com

Rust SIMD on the GPU

Why it matters — Engineers can write a single Rust function using portable SIMD types and have it execute on both CPUs and GPUs without rewriting for vendor intrinsics. The approach treats a GPU warp as a vector unit, so the same arithmetic, comparison, and reduction code maps to a single warp instruction. Adoption requires a Rust toolchain that emits GPU kernels and enables the portable_simd feature, but no special GPU annotations are needed.

2 feeds
9 min
09 258 new

Performance wafer.ai

Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

Why it matters — For teams deploying frontier-scale models, the MI355X's 288GB VRAM enables single-node deployments that would require multi-node setups on B200s, yielding a practical cost advantage at 48 tok/s/$ versus B300's 33 tok/s/$. The tradeoff is significantly slower prefill performance and lingering ROCm software gaps that demand engineering effort, such as a missing top-k renorm function that crashed the speculative decode scheduler.

1 feed
5 min
10 258 new

Performance Techmeme

Meta releases Muse Spark 1.3 in Muse Code and API with claimed coding and agentic gains at unchanged pricing

Why it matters — Teams already on Spark 1.2 get a claimed performance bump at no additional cost, which simplifies upgrade decisions. The emphasis on agentic performance signals Meta is pushing toward multi-step autonomous workflows, though no benchmarks or specifics are provided in the available material. Without independent verification, the significance of the improvements remains unconfirmed.

2 feeds
80 min
11 257 new

Performance github.com

Show HN: Shitty – fast terminal. Memory-unsafe and faster than yours

Why it matters — Engineers who need low-latency, high-throughput terminal I/O can gain measurable speed improvements over popular alternatives. The project’s reliance on specific graphics drivers and a non-standard C++ toolchain limits where it can be deployed, and its memory-unsafe label suggests extra caution for production use.

1 feed
7 min
13 252 new

Performance github.com

Assembly Hall of Shame

Why it matters — Engineers can see which instructions suffer the most from microcode assists, cache-line splits, or uncore traffic, revealing hidden worst-case paths. This insight helps in sizing timing budgets for real-time or safety-critical code and in evaluating the impact of contention-based attacks.

2 feeds
8 min
14 252 new

Performance build2.org

Faster Than Ninja

Why it matters — For engineers choosing a build system, this comparison indicates that Ninja's speed advantage is partly due to offloading work to a generation step (like CMake), which adds time. build2 offers more built-in features (like token-based change tracking) that can be disabled to achieve similar performance, giving teams flexibility without sacrificing speed.

2 feeds
12 min
15 252 new

Performance apple.com

Apple debuts M6 as first 2nm chip and M5 Ultra as first quad-die M-series SoC

Why it matters — M6's 2nm process and Dual Neural Engine deliver up to 2x peak AI compute and nearly 30% more GPU AI performance over M5, making on-device LLM workflows significantly faster. M5 Ultra's quad-die architecture provides 1.2TB/s unified memory bandwidth, 50% more than M3 Ultra, enabling desktop machines to run massive AI models locally.

2 feeds
15 min
17 247 new

Performance Waymo

Waymo reveals custom 5nm ASIC and full-stack compute for fully autonomous driving

Why it matters — Autonomous driving requires deterministic, low-latency compute that off-the-shelf hardware cannot reliably provide. Waymo’s custom silicon and full-stack optimizations demonstrate the scale of investment needed to meet safety and performance demands. This sets a benchmark for edge AI compute in safety-critical applications.

2 feeds
4 min
19 239 new

Performance Linebender

fearless_simd v0.7 adds 64-bit integers, explicit SSE2 level, and improved generics ahead of v1.0

Why it matters — The 64-bit integer support completes full type coverage for integer and float vectors, removing a gap caused by uneven hardware support for 64-bit SIMD operations. The explicit SSE2 level lets crates that don't need runtime dispatch avoid its overhead while using real SIMD intrinsics rather than scalar fallback. Improved trait-based generics make it practical to write functions generic over vector types without resorting to macros or additional crates like paste.

2 feeds
6 min
20 234 new

Performance lemire.me

Profile-guided optimization in Go

Why it matters — For performance-sensitive Go applications, PGO offers a low-effort path to small but measurable gains by replacing compiler heuristics with actual runtime data. The process requires collecting a representative profile and performing a second build, but carries a low risk of significant regressions on unprofiled workloads due to the conservative nature of Go's optimizations.

2 feeds
4 min