llama.cpp C API Guide
Comprehensive reference for the llama.cpp C API, documenting all non-deprecated functions and common usage patterns.
Overview
llama.cpp is a C/C++ implementation for LLM inference with minimal dependencies and state-of-the-art performance. This skill provides:
- Complete API Reference: All non-deprecated functions organized by category
- Common Workflows: Working examples for typical use cases
- Best Practices: Patterns for efficient and correct API usage
Quick Start
See [references/workflows.md](references/workflows.md) for complete working examples. Basic workflow:
llamabackendinit() - Initialize backend
llamamodelloadfromfile() - Load model
llamainitfrom_model() - Create context
llama_tokenize() - Convert text to tokens
llama_decode() - Process tokens
llamasamplersample() - Sample next token
- Cleanup in reverse order
When to Use This Skill
Use this skill when:
- API Lookup: You need to find a specific function (e.g., "How do I load a model?", "What function creates a context?")
- Code Generation: You're writing C code that uses llama.cpp
- Workflow Guidance: You need to understand the steps for a task (e.g., text generation, embeddings, chat)
- Advanced Features: You're working with batches, sequences, LoRA adapters, state management, or custom sampling
- Migration: You're updating code from deprecated functions to current API
Core Concepts
Key Objects
llama_model: Loaded model weights and architecture
llama_context: Inference state (KV cache, compute buffers)
llama_batch: Input tokens and positions for processing
llama_sampler: Token sampling configuration
llama_vocab: Vocabulary and tokenizer
llamamemoryt: KV cache memory handle
Typical Flow
- Initialize:
llamabackendinit()
- Load Model:
llamamodelloadfromfile()
- Create Context:
llamainitfrom_model()
- Tokenize:
llama_tokenize()
- Process:
llamaencode() or llamadecode()
- Sample:
llamasamplersample()
- Generate: Repeat steps 5-6
- Cleanup: Free in reverse order
API Reference
For detailed API documentation, the complete API is split across 6 files for efficient targeted loading. Start with [references/api-core.md](references/api-core.md) which links to all other sections.
API Files:
- [api-core.md](references/api-core.md) (372 lines) - Initialization, parameters, model loading, quantization structs
- [api-model-info.md](references/api-model-info.md) (241 lines) - Model properties, architecture detection, metadata enums
- [api-context.md](references/api-context.md) (423 lines) - Context, memory (KV cache), state management
- [api-inference.md](references/api-inference.md) (420 lines) - Batch operations, inference, tokenization, chat
- [api-sampling.md](references/api-sampling.md) (524 lines) - All 20+ sampling strategies (incl. adaptive-p) + backend sampling API
- [api-advanced.md](references/api-advanced.md) (401 lines) - LoRA adapters, performance, training, constants
Total: 204 active functions (b10665) across 6 organized files
Quick Function Lookup
Most common: llamabackendinit(), llamamodelloadfromfile(), llamainitfrommodel(), llamatokenize(), llamadecode(), llamasamplersample(), llamavocabiseog(), llamamemoryclear()
See [references/api-core.md](references/api-core.md) for the full API index linking to all function signatures.
Common Workflows
See [references/workflows.md](references/workflows.md) for 13 complete working examples: basic text generation, chat, embeddings, batch processing, multi-sequence, LoRA, state save/load, custom sampling (XTC/DRY), encoder-decoder models, model detection, and memory management patterns.
Best Practices
See [references/workflows.md](references/workflows.md) for detailed best practices. Key points:
- Always use default parameter functions (
llamamodeldefault_params(), etc.)
- Check return values for errors
- Free resources in reverse order of creation
- Handle dynamic buffer sizes for tokenization
- Query actual context size after creation (
llamanctx())
- Check for end-of-generation with
llamavocabis_eog()
Common Patterns
End-of-generation check (llamavocabiseog()), logits retrieval (llamagetlogitsith()), batch creation (llamabatchget_one()), tokenization buffer handling. See [references/workflows.md](references/workflows.md) for complete code examples.
Troubleshooting
Common Issues
Model loading fails:
- Verify file path and GGUF format validity
- Check available RAM/VRAM for model size
- Reduce
ngpulayers if GPU memory insufficient
Tokenization returns negative value:
- Buffer too small; reallocate with
-n size and retry
- See tokenization pattern in [Common Patterns](#common-patterns)
Decode/encode returns non-zero:
- Verify batch initialization (
llamabatchgetone() or llamabatch_init())
- Check context capacity (
llamanctx())
- Ensure positions within context window
Silent failures / no output:
- Check if
llamavocabis_eog() immediately returns true
- Verify sampler initialization
- Enable logging:
llamalogset()
Performance issues:
- Increase
n_threads for CPU
- Set
ngpulayers for GPU offloading
- Use larger
n_batch for prompts
- See [Performance & Utilities](references/api-advanced.md#performance--utilities)
Sliding Window Attention (SWA) issues:
- If using Mistral-style models with SWA, set
ctxparams.swafull = true to access beyond attention window
- Check:
llamamodeln_swa(model) to detect SWA size and configuration needs
- Symptoms: Token positions beyond window size causing decode errors
Per-sequence state errors:
- Ensure sequence ID matches when loading:
llamastateseqloadfile(ctx, "file", destseqid, ...)
- Verify token buffer is large enough for loaded tokens
- Check sequence wasn't cleared or removed before loading state
Model type detection:
- Use
llamamodelhas_encoder() before assuming decoder-only architecture
- For recurrent models (Mamba/RWKV), KV cache behavior differs from standard transformers
- Encoder-decoder models require
llamaencode() then llamadecode() workflow
For advanced issues: https://github.com/ggerganov/llama.cpp/discussions
Resources
- API Reference (6 files, 2,354 lines total) - Complete API reference split by category for targeted loading:
- [api-core.md](references/api-core.md) - Initialization, parameters, model loading, quantization structs - [api-model-info.md](references/api-model-info.md) - Model properties, architecture detection, metadata enums - [api-context.md](references/api-context.md) - Context, memory, state management - [api-inference.md](references/api-inference.md) - Batch, inference, tokenization, chat - [api-sampling.md](references/api-sampling.md) - All 20+ sampling strategies (incl. adaptive-p) + backend sampling API - [api-advanced.md](references/api-advanced.md) - LoRA, performance, training, constants
- [references/workflows.md](references/workflows.md) (1,619 lines) - 15 complete working examples: basic workflows (text generation, chat, embeddings, batching, sequences), intermediate (LoRA, state, sampling, encoder-decoder, memory), advanced features (XTC/DRY, per-sequence state, model detection), and production applications (interactive chat, streaming).
What's New in b10665
b10665 (249 commits since b10416) — no functions added, removed, or re-signatured. The public C API changed only in struct fields, one new enum, and the state-file versions:
BREAKING (data, not code):
LLAMASESSIONVERSION 9 → 10 and LLAMASTATESEQVERSION 2 → 3 (recurrent-state rollback in ggmlssmscan changed the serialized layout). Session/sequence-state files from older builds are rejected — llamastateloadfile() / llamastateseqloadfile() fail on them. Regenerate cached sessions; check the return value and fall back to re-ingesting the prompt.
New lazy tensor reading:
enum llamatensorreadlazy (LLAMATENSORREADLAZYOFF/AUTO/ON) + llamamodelparams.tensorread_lazy. Faults in rows of arch-marked tensors on demand instead of reading them whole at load — cuts resident memory and load latency for models with very large sparsely-used tensors. AUTO applies only to marked tensors > 4 GiB; both AUTO and ON require an mmap load mode. See [api-core.md](references/api-core.md#lazy-tensor-reading).
New quantizer memory cap:
llamamodelquantizeparams.maxbufsize (sizet) — max bytes of tensor rows held in memory at once, 0 = default (8 GiB). Lower it to quantize very large models on memory-constrained machines.
Previously (b10416, 158 commits since b10258) — multi-output backend sampling, versioning, and a DRY sampler signature change:
BREAKING:
llamasamplerinitdry() no longer takes the int32t nctxtrain parameter (previously the 2nd argument, right after vocab). Update all call sites — see [api-sampling.md](references/api-sampling.md#dry-sampler).
New multi-output backend sampling [EXPERIMENTAL]:
llamacontextparams.noutputsmaxperseq (uint32t) — max sampled outputs per sequence in a ubatch (0 = noutputs_max).
llamasampleri.backendinit() gained a uint32t noutputsmaxperseq parameter; new backendreset() / copystate() callbacks for authors of custom backend samplers.
llamasamplercopy() — copy mutable sampler state between two same-type/config samplers without losing the destination's bound compute graph.
llamagetsampledtokenith(): with multiple outputs, sampler state now advances on token acceptance, not on read; accept a contiguous prefix in output order (no gaps).
New load-mode default:
LLAMALOADMODEAUTO (-1) is the new default for llamamodelparams.loadmode (was LLAMALOADMODE_MMAP). Auto-detects based on device capabilities — e.g. avoids mmap on iGPUs.
New:
llama_version() — returns the llama.cpp library version string (from CMake's new semantic-versioning setup).
Doc correction (no code change): llamasamplerinitpenalties()'s penaltylastn doc previously said "-1 = context size"; that was never accurate — negative values are clamped to 0 (disabled). History-based samplers (DRY, penalties) no longer resolve a "full-context window" from training context size, since backend sampling constructs samplers before a llamacontext (and its resolved context length) exists.
Not in llama.h/llama-cpp.h, mentioned for awareness: mtmd gained Qwen3-TTS support — a breaking change to the llama-tts CLI binary (outside this skill's C API scope).
Previously (b10258, 183 commits since b10075):
BREAKING:
llamasamplerinitpenalties() gained a new required first parameter nvocab (source it via llamavocabn_tokens(vocab)). Update all call sites.
New load-mode API (replaces three booleans):
enum llamaloadmode (LLAMALOADMODENONE/MMAP/MLOCK/MMAPMLOCK/DIRECTIO) + llamaloadmodename() / llamaloadmodefromstr().
llamamodelparams.loadmode replaces the removed usemmap, usedirectio, and use_mlock boolean fields.
llamamodelparams.load_mtp (bool) — whether to load MTP layers.
New vocab function:
llamavocabgetsuppresstokens() — model-specific suppress tokens (gguf key tokenizer.ggml.suppress_tokens).
Previously (b10075, ~205 commits since b9870): touched the public C API in exactly one place — new LLAMAFTYPEMOSTLYQ20 = 41 quantization type (CPU backend), PR #24448. Everything else in that range (new Hy3/hyv3 model + MTP speculative decoding, server reasoningbudget_tokens, mtmd NUL-truncation fix, deepseek-ocr v1 multi-tile, /responses streaming timings, etc.) was internal/server-side and did not change llama.h.
Recent (added in b9859→b9870, PR #25134):
llamamodelftype() — returns the model's file type as an enum llamaftype (e.g. LLAMAFTYPEMOSTLYQ8_0).
llamaftypename() — converts an enum llamaftype to a human-readable string (e.g. "Q80", "Q4_K - Medium"). Pair the two to display a loaded model's quantization.
Recent (added in b9840, from b9704):
llamamodelnlayernextn() — returns the number of NextN (Multi-Token Prediction / MTP) layers in the model. These speculative next-token prediction layers power MTP-capable architectures such as DeepSeek V3/V4, GLM, Qwen3.5-MoE, and Step3.5. Returns 0 for non-MTP models. The total layer count equals llamamodeln_layer() (effective layers) + this value.
Recent (added in b9704) [EXPERIMENTAL]:
llamacontextparams.noutputsmax (uint32t) — max outputs in a ubatch (0 = nbatch). Cap it to reserve less output VRAM when you read only a few logits/embeddings per batch (e.g. one output per sequence during generation).
llamacontextparams.ctxother (struct llamacontext *) — a source/target/parent context for sharing inference results or llama_memory (KV cache) between two contexts; used by MTP setups such as Gemma4 MTP.
llamasetwarmup() deprecated — perform warmup runs manually instead. It changed graph topology with MoE models (causing extra reallocations) and will be removed in a future release.
Still recent (added in b9246) [EXPERIMENTAL]:
- Multi-Token Prediction:
enum llamacontexttype (LLAMACONTEXTTYPEDEFAULT/LLAMACONTEXTTYPEMTP) + llamacontextparams.ctx_type
- Recurrent-state rollback (Mamba/RWKV):
llamacontextparams.nrsseq + llamanrs_seq(ctx)
- Sequence-state flags:
LLAMASTATESEQFLAGSNONE (0), LLAMASTATESEQFLAGSON_DEVICE (2)
Stable Since b8809:
- Model loading:
llamamodelloadfromfileptr(), llamamodelinitfrom_user()
- Quantization types: MXFP4MOE (38), NVFP4 (39), Q10 (40)
- Split mode:
LLAMASPLITMODE_TENSOR (3) for backend-agnostic tensor parallelism
- Backend sampling API (EXPERIMENTAL): GPU-accelerated sampling via context params
- Adaptive-P sampler:
llamasamplerinitadaptivep()
Key Differences from Deprecated API
If you're updating old code:
- Use
llamamodelloadfromfile() instead of llamaloadmodelfromfile()
- Use
llamamodelfree() instead of llamafreemodel()
- Use
llamainitfrommodel() instead of llamanewcontextwith_model()
- Use
llamavocab() functions instead of llamatoken()
- Use
llamastate*() functions instead of deprecated state functions
- Use
llamasetadapterslora() instead of llamasetadapterlora() for LoRA adapters
- Use
llamavocabbos() instead of llamavocabcls() (CLS is equivalent to BOS)
- Use
llamasamplerinitgrammarlazypatterns() instead of llamasamplerinitgrammar_lazy()
- Perform warmup runs manually instead of calling deprecated
llamasetwarmup() (deprecated in b9704)
See the API reference for complete mappings.