Oscar Alateras

I work on LLM inference — how models actually run, and what limits them.

I build things and measure them — serving systems, kernels, benchmarks — and publish what I find, including when the result contradicts what I expected. Currently an AI Architect at tne.ai, finishing a Masters in Computer Science.

Projects

inference-arch-fundamentals

Working through the hardware that decides how fast a model can be served, from first principles. Each topic gets its own benchmark, measured results and write-up — not a summary of what the textbook says, but what my own machine actually did. 3 of 13 topics complete.

More on what I'm building