Speculative Decoding — Naming Conventions
Apply this skill when adding, renaming, or reviewing identifiers in speculative decoding code (anything under python/sglang/srt/speculative/, related attention backends, scheduler accumulators, IPC fields, observability metrics, or CLI flags).
Rule 1 — Verb form, drop -ed
Use the verb form accept everywhere. Don't use the past-participle form accepted.
| Don't |
Do |
numacceptedtokens |
numaccepttokens |
accepted_indices |
accept_indices |
acceptedtokenids |
accept_tokens (also see Rule 3) |
Rule 2 — The extra/bonus token is bonustoken / bonustokens
The "+1" token that the target model always emits in addition to verifying drafts is the bonus token. Use bonustoken / bonustokens per Rule 7.
| Don't |
Do |
verifiedid / verifiedids |
bonustoken / bonustokens |
outputid / outputids (when referring to the bonus) |
bonustoken / bonustokens |
req.output_ids (the full output history of a request) is unrelated and stays as is.
Rule 3 — accept includes bonus; correct excludes bonus
The semantic distinction lives in the verb, not the noun. Don't enumerate noun pairs.
| Verb |
Meaning |
**accept_*** |
Includes the bonus token |
**correct_*** |
Drafts only, no bonus |
Pair with whatever noun fits the data (tokens, drafts, indices, …). No required pairing, but preferred default nouns: accepttokens and correctdrafts — correct semantically describes drafts (what got verified), accept describes the resulting token sequence (incl. bonus).
| Form |
Meaning |
accepttokens / acceptindices |
Include bonus |
correct_drafts |
Drafts only, no bonus |
numaccepttokens |
Count incl. bonus |
numcorrectdrafts |
Count excl. bonus |
Exception: acceptrate / acceptlength follow paper convention
These two metric names are entrenched in the spec-decoding literature and in external-facing fields (meta_info, Prometheus). Their semantics are paper-defined, not Rule-3-defined:
| Name |
Paper term |
Bonus? |
Definition |
accept_rate |
$\alpha$ (Leviathan 2023) |
No |
per-draft-token acceptance probability = correctdrafts / proposeddrafts |
accept_length |
$\tau$ (EAGLE) |
Yes |
avg tokens per verify step = completiontokens / verifyct |
Internal counters still follow Rule 3 strict semantics: numcorrectdrafts (no bonus), numaccepttokens (with bonus).
Rule 4 — num for counts; ct for counters; _rate for rates; no prefix for IDs
Each form has its own marker. Never mix (no numXct, no numacceptrate).
| Form |
Pattern |
Meaning |
Examples |
| Count |
num_X |
Snapshot quantity at one point in time (often a tensor or scalar) |
numaccepttokens, numcorrectdrafts, numproposeddrafts |
| Counter |
X_ct |
Monotonically incrementing accumulator over time |
specverifyct, forward_ct |
| Rate / ratio |
X_rate |
Fractional value in [0, 1] |
accept_rate |
| Tokens / content array |
no prefix |
The actual token data, not a count |
accepttokens, correctdrafts, bonus_token |
Rule 5 — Drop redundant tokenid / tokenids suffix in spec scope
id / ids and token / tokens are both fine. But don't combine — tokenid / tokenids is redundant inside spec decoding, because spec code only ever deals with vocab integers.
The semantic differs by scope:
| Scope |
Example |
What tokenid means |
| Framework / multimodal / tokenizer |
imagetokenid, padtokenid, eostokenid, masktokenid, bostokenid |
A specific named/role token's vocab ID. The prefix names the role; tokenid says it's the integer ID for that role. Both halves carry information. |
| Spec decoding |
acceptedtokenids, currtokenid, outtokenids |
Redundant. Spec only deals with vocab integers; id adds nothing beyond token. |
Renames
| Don't |
Do |
acceptedtokenids |
accept_tokens (Rule 1 + 3) |
currtokenid |
current_token |
outtokenids |
out_tokens |
resolvespecoverlaptoken_ids |
resolvespecoverlaptokens |
Rule 6 — Singular vs plural
Plural for any non-scalar tensor ([bs]-shaped, flat, or multi-dim); singular only for scalars (kernel tl.load results, single-int locals). Applies to all spec-decoding tensors (tokens, indices, etc.).
accept_tokens: torch.Tensor # [total_accepted] flat - plural
accept_indices: torch.Tensor # [bs, num_draft_tokens] - plural
draft_tokens: torch.Tensor # [bs * num_draft_tokens] flat - plural
bonus_tokens: torch.Tensor # [bs] - plural
accept_token = tl.load(...) # int32 scalar in a kernel iteration - singular
bonus_token = tl.load(...) # int32 scalar inside a kernel - singular
Out of scope (these names stay as is)
These rules apply to spec-decoding-specific identifiers. Pre-existing or framework-level names are kept.
- PyTorch / ecosystem:
seqlens, extendseqlens, cuseqlens_q
- Framework / multimodal vocab:
imagetokenid, padtokenid, eostokenid, masktokenid, hottokenid, bostokenid, topk_id
- Request-level state:
req.inputids, req.outputids, req.origininputids, nexttokenids (model_runner.sample output)
- Frozen C++ kwargs:
accepttokennum (sgl-kernel)
- Non-token IDs:
reqid, gpuid, layerid, programid
len / lens names: numX is preferred for counts (Rule 4), but len / lens names are acceptable. Triton kernel params in particular often use lens / len to align with the PyTorch ecosystem (seqlens, cuseqlensq). Rule 1 still requires the -ed-less form (acceptlength OK, acceptedlength not).