🌱

Green AI Index

The world's first open leaderboard for AI energy efficiency.

kWh per 1M tokens · Community-sourced · Anonymous · ≥10 submissions per group

EU AI ActGHG ProtocolSB 253

Submit to the leaderboard

nemulai benchmark --model-tag llama-3-8b --throughput 1500 --upload

What is throughput?

Technical reports

TR-2026-03July 2026 · Azure production traces · simulation

Margin Dispersion Under Proxy-Metric Billing: A Trace-Driven Simulation of Multi-Tenant LLM Inference Economics

Token pricing is the tightest billing proxy there is — and it still leaks. If tenants differ in prefill:decode mix, identical prices produce structural margin dispersion, widening sharply as blended margin compresses. Conservation-checked attribution, full source published.

TR-2026-02July 2026 · H100 · vLLM

Whole-GPU Idle Under Bursty Request Arrival: Measuring the Gap Between Billed Instance-Hours and Consumed GPU-Seconds on an LLM Serving Endpoint

GPU-busy falls from 99.9% to 2.0% as request arrival goes sparse — even a 1 req/s endpoint idles 69% of its billed window. Five controlled arrival patterns, pre-registered prediction, raw NVML streams published.

Loading…

All submissions are anonymized using HMAC-SHA256. Individual users are never identified. Groups with fewer than 10 submissions are excluded.

Enable benchmarking in your dashboard to contribute your data.