block #0001 in --:--:--Join the pool
PENDING…
ai pending 5/6 1h ago · 3 min read

Magnitude inference tunes kernels on your box

Magnitude inference launched on HN as a local engine that autotunes kernels per machine. HN and GIGAZINE cover the 2x-vs-llama.cpp claim with caveats.

pending 5/6 — still in the mempool

Early story. Some claims here are not officially confirmed yet. We update this post as it confirms.

Magnitude inference tunes kernels on your box
tl;dr
  • The Launch HN post by anerli and GIGAZINE's October 1 note both describe Magnitude as a self-optimizing inference engine for agents with Apple Silicon, NVIDIA, and AMD paths.
  • Both carry the headline speed claim against llama.cpp, including Metal numbers on an M4 Pro class machine.
  • HN comments immediately poke holes on MLX comparisons and M5 gaps; those caveats stay in the story.
in this block
  1. What actually happened
  2. Benches that do not travel alone
  3. Agent-shaped constraints
  4. How to read the Mac fight
  5. What to do as a reader

Magnitude inference showed up on Hacker News as a YC S25 Launch HN about a day before this desk wrote it up, and GIGAZINE's October 1 English digest covers the same engine: an open-source local stack that profiles your machine, tunes kernels on-device, and aims at agent sessions rather than datacenter batching.

What actually happened

Magnitude inference is not another chat UI. The founders say they previously shipped a browser agent, then needed local models that did not fall over on long, concurrent agent sessions. Their Launch HN post frames three competitor tradeoffs: datacenter batch engines, broad-compatibility engines like llama.cpp, and narrow hardware specialists. Magnitude's pitch is on-device compilation and tuning, kernels focused on popular open-weight families, dynamic memory that grows with sessions, and hybrid paged attention for prefix sharing without wrecking single-session speed.

License and stack details are in the HN post: Apache 2.0, Rust, custom GPU kernel runtime and autotuner. GIGAZINE summarizes the same compatibility list and the "up to 2x" framing. Shared launch facts stay shared. Deep kernel theology stays with the founders' thread.

Benches that do not travel alone

On Metal with Qwen 3.6 35B A3B at 4-bit and 64k context, the HN post prints 30 tok/s to 57 tok/s decode versus llama.cpp, 466 to 507 tok/s prefill, and 28% less per-agent memory. On CUDA with a DGX Spark class box, it prints 49 to 58 tok/s decode and 2,033 to 2,507 tok/s prefill. GIGAZINE repeats the M4 Pro style comparison. Commenters on HN say llama.cpp is a low bar on Mac versus MLX, that some M5 machines saw Magnitude slower, and that multi-GPU detection still miscounts cards. The founders reply that M5 Metal 4 matmul paths are a known gap. memcool does not flatten those replies into "2x everywhere."

Agent-shaped constraints

The HN post stresses long sessions, several at once, and still using the machine for other work. Dynamic memory that only reserves weights up front is aimed at that. Speculative decoding drafters, TurboQuant-inspired KV quantization to 8-bit keys and 4-bit values, and a desktop app that wakes models for Pi, OpenCode, Hermes, or Codex are part of the same product story. Those features are founder claims. Commenters asking for ROCm, multi-machine sharding, or Whisper are feature requests, not shipped checkmarks.

Business model chatter on the thread points to a future hybrid cloud. That is aspiration. The artifact in front of readers today is the local engine and the open repo.

How to read the Mac fight

Several HN commenters say the right Apple baseline is MLX, not llama.cpp. The founders say rough MLX comparisons still favor Magnitude on decode in their internal checks and promise a public MLX bench. Until that lands, treat Metal-versus-llama.cpp charts as directional, not as a death match against every MLX build. Another commenter posted slower Magnitude numbers on a fresh llama.cpp build; the founders again pointed at M5 matmul gaps. That back-and-forth is why this desk refuses a single 2x slogan without hardware tags.

Prefix-cache reuse across agent turns is the product question that matters more than a short-context tok/s screenshot. The founders say a prefix tree reuses prompts across sessions and that the HTTP surface is OpenAI-compatible chat completions. If your harness already speaks llama.cpp OpenAI wrapper, that is the integration path they are selling.

What to do as a reader

If you run local agents, Magnitude inference is worth a weekend bake-off against your current llama.cpp or MLX path on your actual hardware, not on their screenshot. Read the Launch HN thread and the GIGAZINE note. Compare also with OpenClaw Enterprise if you care about control planes more than kernels, or Robinhood Agents if your "agent" is a broker feature. Nothing here says uninstall llama.cpp tonight. Not financial advice. DYOR, ser.

If the desktop app spends too long assessing models before downloads, that is a known friction called out in the HN thread, not a feature.

Not financial advice. DYOR, ser.

More in the pool

all ai
gm ser

Get confirmed before the crowd

Daily block at 07:00 UTC. No spam, just the block, ser.