Inference Engineering
Philip Kiely · Baseten
“The more constraints you can introduce into your inference system, the better performance you will achieve.” The Golden Rule of Inference.
Tinkering with AI from a Linux box in Paris.
I'm Keylhan, an engineering student at EFREI Paris (just finished my first year, année préparatoire), French, daily-driving Linux with Windows on the side for school. I read a lot, lurk on X more than I post, and tinker with AI the way the hardware allows: no GPUs, so I go where the interesting work doesn't need them, going after reverse-engineering how tools are built, local-first infra, clever scripts, and the odd borderline-legal extraction.
3 read · 2 on the queue
Philip Kiely · Baseten
“The more constraints you can introduce into your inference system, the better performance you will achieve.” The Golden Rule of Inference.
Andrew Hunt & David Thomas · 20th Anniversary Edition
Good design is always easier to change, and great code starts with taking ownership.
Italo Calvino
“A classic is a book that has never finished saying what it has to say.” Fourteen definitions of a classic, and no reason better than this one.
Malcolm Gladwell
“The right way to talk to strangers is with caution and humility.” The closing thesis of a book about how badly we misread people we don't know.
ongoing, broad goal
Read as much as I can from O'Reilly over the coming months / years.
6 entries · no GPUs harmed
Reverse-engineered how Claude's web sandbox works and rebuilt it from the ground up.
Curiosity-driven teardown. Figured out the moving parts, then re-implemented them myself to actually understand the design.
A script that skips the 20-minute llama.cpp compile on 4-core Colab sessions.
I burned a lot of Colab sessions recompiling llama.cpp every time: 4 cores, 20+ minutes each. So I scripted the painful part away. Now a fresh cloud instance is usable in minutes.
Extracted the .exe, rebuilt the source tree from the extracted contents, now I can recompile with my own patches.
A bit borderline, legally, but a great exercise. Currently working on recovering meaningful function names from the minified output, so LLMs can actually navigate the codebase instead of drowning in garbage symbols.
Security guards, disciplined multi-agent orchestration, and persistent memory, for Claude Code and opencode.
Zenno watches what agents do and keeps them honest: four Shield guards block secret leaks, confidential-file reads, risky installs, and self-modification of Zenno itself; hard caps stop runaway subagent fan-out; a memory layer carries project knowledge across sessions; and a doctor, cost tracker, and trace exporter keep the whole thing observable. I run it on my own machine every day.
A voice-controlled clock that runs its command routing on-device: mic → Voxtral STT → a locally fine-tuned Laya router, cloud LLM only as fallback. Speaks French and English, including mixed.
The interesting part wasn't the app, it was the measurement. Zero-shot routing scored 91% in English but 62% in French, so I generated 2,500 labeled commands, fine-tuned Laya's multilingual checkpoint with RLCD on a Colab T4, and took French to 93%, then spent two more rounds chasing the confident mistakes that remained, and learned when to stop training and fix safety in plain code instead.
13 compose projects on an OptiPlex-3070, all behind Traefik, plus a Hermes agent that does the daily driving.
The substrate under most of my experiments: photos, media, bookmarks, DNS, and the dashboards that tell me the box is still alive. Intel UHD 630, no discrete GPU, everything on one 1TB drive and a 256GB NVMe. Watchtower keeps the images fresh, which is the least glamorous part of owning any of this.
Aug 1 – Sep 28, 2026 across opencode + Claude Code, 283 sessions, 5.04B tokens processed, $0.00 spent. Being a free-tier hacker has its perks. Sessions count the ones I started myself, with subagent fan-out left out.
exported from opencode + Claude Code analytics · Aug 01, 2026 → Sep 28, 2026 · updated Sep 28, 2026. cache reads account for 92% of all tokens: the prompt cache does the heavy lifting. analytics tools estimate $1,891 for the same usage, since they price free models at list rate.
I lurk more than I post, but I read everything and I'm always up for a conversation about AI tooling, reverse-engineering, or a good book.