AI inference efficiency

Faster inference, proven before it is claimed.

Effiq is an AI compute-efficiency company. We optimize large-model inference — and every performance statement we publish comes with the locked protocol and the raw artifacts needed to re-run it, disagree with it, or verify it independently.

Read the method↗ See a live run↗
What we do

Optimization is the product. Measurement is the moat.

Anyone can report a speedup. We ship the proof with it.

LLM inference economics are won or lost on details: quantization strategy, batching and scheduling, memory layout, kernel choice. We build and apply those optimizations where they measurably hold up — on your hardware, your workload, your serving stack.

What makes an Effiq number different is not the number. It is that the comparison was locked in writing before the data existed, executed against that lock, and decided by a statistical rule that cannot be bargained with after the fact. When the evidence does not clear the bar, we publish that too.

01

Lock the rules first

Hypothesis, hardware, workload, metric, decision threshold, and failure handling are committed in a public pre-registration before any measurement. Changing them later requires a dated, written deviation — and the change itself becomes part of the record.

02

Execute and preserve

Runs are interleaved, servers restarted between arms, and every raw artifact kept: per-request records, artifact hashes checked against publisher records, environment versions, and a timestamped execution log.

03

Decide by the bound, not the vibe

A claim passes only if the lower bound of the confidence interval on the speedup ratio clears the pre-registered threshold. Point estimates do not win arguments. The full methodology is public in the protocol repository.

Public evidence

Everything we claim in public, you can re-run in private.

github.com/effiq/protocol Methodology · published before any numbers

Open templates for trustworthy inference benchmarks: pre-registration, decision rules, execution logs, one-day re-run checklists, and a deviation procedure. We published our method first and accepted the risk of being copied — because the discipline, not the secret, is the advantage. License: Apache-2.0 · Status: maintained

github.com/effiq/practice Run 001 · pre-registered, executed, reported

A complete end-to-end run at toy scale: two weight artifacts of the same model compared on identical serving settings, six interleaved runs, bootstrap confidence interval, quality gate. The candidate's speedup was not confirmed under the locked rule — and we published that result anyway, along with an artifact-integrity deviation caught by hash verification before any data was collected. Decision rule: FAIL — R = 0.909, 95% CI [0.629, 1.301], pass bar ≥ 1.20 · Quality gate: PASS (0 pp degradation)