smithery.ai

llamacpp

Complete llama.cpp C/C++ API reference covering model loading, inference, text generation, embeddings, chat, tokenization, sampling, batching, KV cache, LoRA adapters, and state management.

First seen Mar 28, 2026

Installation

$ npx skills add https://smithery.ai

Summary

  • Complete llama.cpp C/C++ API reference covering model loading, inference, text generation, embeddings, chat, tokenization, sampling, batching, KV cache, LoRA adapters, and state management.
  • Triggers on: llama.cpp questions, LLM inference code, GGUF models, local AI/ML inference, C/C++ LLM integration, \"how do I use llama.cpp\", API function lookups, implementation questions, troubleshooting llama.cpp issues, and any llama-cpp or ggerganov/llama.cpp mentions.

Also in this package

Other skills from smithery.ai · top by installs.

npx skills add https://smithery.ai

Browse all from smithery.ai

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 15,662 B
  • docs SUMMARY.md 477 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 1 installs

SKILL.md

llama.cpp C API Guide

Comprehensive reference for the llama.cpp C API, documenting all non-deprecated functions and common usage patterns.

Overview

llama.cpp is a C/C++ implementation for LLM inference with minimal dependencies and state-of-the-art performance. This skill provides:

  • Complete API Reference: All non-deprecated functions organized by category
  • Common Workflows: Working examples for typical use cases
  • Best Practices: Patterns for efficient and correct API usage

Quick Start

See [references/workflows.md](references/workflows.md) for complete working examples. Basic workflow:

  1. llamabackendinit() - Initialize backend
  2. llamamodelloadfromfile() - Load model
  3. llamainitfrom_model() - Create context
  4. llama_tokenize() - Convert text to tokens
  5. llama_decode() - Process tokens
  6. llamasamplersample() - Sample next token
  7. Cleanup in reverse order

When to Use This Skill

Use this skill when:

  1. API Lookup: You need to find a specific function (e.g., "How do I load a model?", "What function creates a context?")
  2. Code Generation: You're writing C code that uses llama.cpp
  3. Workflow Guidance: You need to understand the steps for a task (e.g., text generation, embeddings, chat)
  4. Advanced Features: You're working with batches, sequences, LoRA adapters, state management, or custom sampling
  5. Migration: You're updating code from deprecated functions to current API

Core Concepts

Key Objects

  • llama_model: Loaded model weights and architecture
  • llama_context: Inference state (KV cache, compute buffers)
  • llama_batch: Input tokens and positions for processing
  • llama_sampler: Token sampling configuration
  • llama_vocab: Vocabulary and tokenizer
  • llamamemoryt: KV cache memory handle

Typical Flow

  1. Initialize: llamabackendinit()
  2. Load Model: llamamodelloadfromfile()
  3. Create Context: llamainitfrom_model()
  4. Tokenize: llama_tokenize()
  5. Process: llamaencode() or llamadecode()
  6. Sample: llamasamplersample()
  7. Generate: Repeat steps 5-6
  8. Cleanup: Free in reverse order

API Reference

For detailed API documentation, the complete API is split across 6 files for efficient targeted loading. Start with [references/api-core.md](references/api-core.md) which links to all other sections.

API Files:

  • [api-core.md](references/api-core.md) (372 lines) - Initialization, parameters, model loading, quantization structs
  • [api-model-info.md](references/api-model-info.md) (241 lines) - Model properties, architecture detection, metadata enums
  • [api-context.md](references/api-context.md) (423 lines) - Context, memory (KV cache), state management
  • [api-inference.md](references/api-inference.md) (420 lines) - Batch operations, inference, tokenization, chat
  • [api-sampling.md](references/api-sampling.md) (524 lines) - All 20+ sampling strategies (incl. adaptive-p) + backend sampling API
  • [api-advanced.md](references/api-advanced.md) (401 lines) - LoRA adapters, performance, training, constants

Total: 204 active functions (b10665) across 6 organized files

Quick Function Lookup

Most common: llamabackendinit(), llamamodelloadfromfile(), llamainitfrommodel(), llamatokenize(), llamadecode(), llamasamplersample(), llamavocabiseog(), llamamemoryclear()

See [references/api-core.md](references/api-core.md) for the full API index linking to all function signatures.

Common Workflows

See [references/workflows.md](references/workflows.md) for 13 complete working examples: basic text generation, chat, embeddings, batch processing, multi-sequence, LoRA, state save/load, custom sampling (XTC/DRY), encoder-decoder models, model detection, and memory management patterns.

Best Practices

See [references/workflows.md](references/workflows.md) for detailed best practices. Key points:

  • Always use default parameter functions (llamamodeldefault_params(), etc.)
  • Check return values for errors
  • Free resources in reverse order of creation
  • Handle dynamic buffer sizes for tokenization
  • Query actual context size after creation (llamanctx())
  • Check for end-of-generation with llamavocabis_eog()

Common Patterns

End-of-generation check (llamavocabiseog()), logits retrieval (llamagetlogitsith()), batch creation (llamabatchget_one()), tokenization buffer handling. See [references/workflows.md](references/workflows.md) for complete code examples.

Troubleshooting

Common Issues

Model loading fails:

  • Verify file path and GGUF format validity
  • Check available RAM/VRAM for model size
  • Reduce ngpulayers if GPU memory insufficient

Tokenization returns negative value:

  • Buffer too small; reallocate with -n size and retry
  • See tokenization pattern in [Common Patterns](#common-patterns)

Decode/encode returns non-zero:

  • Verify batch initialization (llamabatchgetone() or llamabatch_init())
  • Check context capacity (llamanctx())
  • Ensure positions within context window

Silent failures / no output:

  • Check if llamavocabis_eog() immediately returns true
  • Verify sampler initialization
  • Enable logging: llamalogset()

Performance issues:

  • Increase n_threads for CPU
  • Set ngpulayers for GPU offloading
  • Use larger n_batch for prompts
  • See [Performance & Utilities](references/api-advanced.md#performance--utilities)

Sliding Window Attention (SWA) issues:

  • If using Mistral-style models with SWA, set ctxparams.swafull = true to access beyond attention window
  • Check: llamamodeln_swa(model) to detect SWA size and configuration needs
  • Symptoms: Token positions beyond window size causing decode errors

Per-sequence state errors:

  • Ensure sequence ID matches when loading: llamastateseqloadfile(ctx, "file", destseqid, ...)
  • Verify token buffer is large enough for loaded tokens
  • Check sequence wasn't cleared or removed before loading state

Model type detection:

  • Use llamamodelhas_encoder() before assuming decoder-only architecture
  • For recurrent models (Mamba/RWKV), KV cache behavior differs from standard transformers
  • Encoder-decoder models require llamaencode() then llamadecode() workflow

For advanced issues: https://github.com/ggerganov/llama.cpp/discussions

Resources

  • API Reference (6 files, 2,354 lines total) - Complete API reference split by category for targeted loading:

- [api-core.md](references/api-core.md) - Initialization, parameters, model loading, quantization structs - [api-model-info.md](references/api-model-info.md) - Model properties, architecture detection, metadata enums - [api-context.md](references/api-context.md) - Context, memory, state management - [api-inference.md](references/api-inference.md) - Batch, inference, tokenization, chat - [api-sampling.md](references/api-sampling.md) - All 20+ sampling strategies (incl. adaptive-p) + backend sampling API - [api-advanced.md](references/api-advanced.md) - LoRA, performance, training, constants

  • [references/workflows.md](references/workflows.md) (1,619 lines) - 15 complete working examples: basic workflows (text generation, chat, embeddings, batching, sequences), intermediate (LoRA, state, sampling, encoder-decoder, memory), advanced features (XTC/DRY, per-sequence state, model detection), and production applications (interactive chat, streaming).

What's New in b10665

b10665 (249 commits since b10416) — no functions added, removed, or re-signatured. The public C API changed only in struct fields, one new enum, and the state-file versions:

BREAKING (data, not code):

  • LLAMASESSIONVERSION 9 → 10 and LLAMASTATESEQVERSION 2 → 3 (recurrent-state rollback in ggmlssmscan changed the serialized layout). Session/sequence-state files from older builds are rejected — llamastateloadfile() / llamastateseqloadfile() fail on them. Regenerate cached sessions; check the return value and fall back to re-ingesting the prompt.

New lazy tensor reading:

  • enum llamatensorreadlazy (LLAMATENSORREADLAZYOFF/AUTO/ON) + llamamodelparams.tensorread_lazy. Faults in rows of arch-marked tensors on demand instead of reading them whole at load — cuts resident memory and load latency for models with very large sparsely-used tensors. AUTO applies only to marked tensors > 4 GiB; both AUTO and ON require an mmap load mode. See [api-core.md](references/api-core.md#lazy-tensor-reading).

New quantizer memory cap:

  • llamamodelquantizeparams.maxbufsize (sizet) — max bytes of tensor rows held in memory at once, 0 = default (8 GiB). Lower it to quantize very large models on memory-constrained machines.

Previously (b10416, 158 commits since b10258) — multi-output backend sampling, versioning, and a DRY sampler signature change:

BREAKING:

  • llamasamplerinitdry() no longer takes the int32t nctxtrain parameter (previously the 2nd argument, right after vocab). Update all call sites — see [api-sampling.md](references/api-sampling.md#dry-sampler).

New multi-output backend sampling [EXPERIMENTAL]:

  • llamacontextparams.noutputsmaxperseq (uint32t) — max sampled outputs per sequence in a ubatch (0 = noutputs_max).
  • llamasampleri.backendinit() gained a uint32t noutputsmaxperseq parameter; new backendreset() / copystate() callbacks for authors of custom backend samplers.
  • llamasamplercopy() — copy mutable sampler state between two same-type/config samplers without losing the destination's bound compute graph.
  • llamagetsampledtokenith(): with multiple outputs, sampler state now advances on token acceptance, not on read; accept a contiguous prefix in output order (no gaps).

New load-mode default:

  • LLAMALOADMODEAUTO (-1) is the new default for llamamodelparams.loadmode (was LLAMALOADMODE_MMAP). Auto-detects based on device capabilities — e.g. avoids mmap on iGPUs.

New:

  • llama_version() — returns the llama.cpp library version string (from CMake's new semantic-versioning setup).

Doc correction (no code change): llamasamplerinitpenalties()'s penaltylastn doc previously said "-1 = context size"; that was never accurate — negative values are clamped to 0 (disabled). History-based samplers (DRY, penalties) no longer resolve a "full-context window" from training context size, since backend sampling constructs samplers before a llamacontext (and its resolved context length) exists.

Not in llama.h/llama-cpp.h, mentioned for awareness: mtmd gained Qwen3-TTS support — a breaking change to the llama-tts CLI binary (outside this skill's C API scope).

Previously (b10258, 183 commits since b10075):

BREAKING:

  • llamasamplerinitpenalties() gained a new required first parameter nvocab (source it via llamavocabn_tokens(vocab)). Update all call sites.

New load-mode API (replaces three booleans):

  • enum llamaloadmode (LLAMALOADMODENONE/MMAP/MLOCK/MMAPMLOCK/DIRECTIO) + llamaloadmodename() / llamaloadmodefromstr().
  • llamamodelparams.loadmode replaces the removed usemmap, usedirectio, and use_mlock boolean fields.
  • llamamodelparams.load_mtp (bool) — whether to load MTP layers.

New vocab function:

  • llamavocabgetsuppresstokens() — model-specific suppress tokens (gguf key tokenizer.ggml.suppress_tokens).

Previously (b10075, ~205 commits since b9870): touched the public C API in exactly one place — new LLAMAFTYPEMOSTLYQ20 = 41 quantization type (CPU backend), PR #24448. Everything else in that range (new Hy3/hyv3 model + MTP speculative decoding, server reasoningbudget_tokens, mtmd NUL-truncation fix, deepseek-ocr v1 multi-tile, /responses streaming timings, etc.) was internal/server-side and did not change llama.h.

Recent (added in b9859→b9870, PR #25134):

  • llamamodelftype() — returns the model's file type as an enum llamaftype (e.g. LLAMAFTYPEMOSTLYQ8_0).
  • llamaftypename() — converts an enum llamaftype to a human-readable string (e.g. "Q80", "Q4_K - Medium"). Pair the two to display a loaded model's quantization.

Recent (added in b9840, from b9704):

  • llamamodelnlayernextn() — returns the number of NextN (Multi-Token Prediction / MTP) layers in the model. These speculative next-token prediction layers power MTP-capable architectures such as DeepSeek V3/V4, GLM, Qwen3.5-MoE, and Step3.5. Returns 0 for non-MTP models. The total layer count equals llamamodeln_layer() (effective layers) + this value.

Recent (added in b9704) [EXPERIMENTAL]:

  • llamacontextparams.noutputsmax (uint32t) — max outputs in a ubatch (0 = nbatch). Cap it to reserve less output VRAM when you read only a few logits/embeddings per batch (e.g. one output per sequence during generation).
  • llamacontextparams.ctxother (struct llamacontext *) — a source/target/parent context for sharing inference results or llama_memory (KV cache) between two contexts; used by MTP setups such as Gemma4 MTP.
  • llamasetwarmup() deprecated — perform warmup runs manually instead. It changed graph topology with MoE models (causing extra reallocations) and will be removed in a future release.

Still recent (added in b9246) [EXPERIMENTAL]:

  • Multi-Token Prediction: enum llamacontexttype (LLAMACONTEXTTYPEDEFAULT/LLAMACONTEXTTYPEMTP) + llamacontextparams.ctx_type
  • Recurrent-state rollback (Mamba/RWKV): llamacontextparams.nrsseq + llamanrs_seq(ctx)
  • Sequence-state flags: LLAMASTATESEQFLAGSNONE (0), LLAMASTATESEQFLAGSON_DEVICE (2)

Stable Since b8809:

  • Model loading: llamamodelloadfromfileptr(), llamamodelinitfrom_user()
  • Quantization types: MXFP4MOE (38), NVFP4 (39), Q10 (40)
  • Split mode: LLAMASPLITMODE_TENSOR (3) for backend-agnostic tensor parallelism
  • Backend sampling API (EXPERIMENTAL): GPU-accelerated sampling via context params
  • Adaptive-P sampler: llamasamplerinitadaptivep()

Key Differences from Deprecated API

If you're updating old code:

  • Use llamamodelloadfromfile() instead of llamaloadmodelfromfile()
  • Use llamamodelfree() instead of llamafreemodel()
  • Use llamainitfrommodel() instead of llamanewcontextwith_model()
  • Use llamavocab() functions instead of llamatoken()
  • Use llamastate*() functions instead of deprecated state functions
  • Use llamasetadapterslora() instead of llamasetadapterlora() for LoRA adapters
  • Use llamavocabbos() instead of llamavocabcls() (CLS is equivalent to BOS)
  • Use llamasamplerinitgrammarlazypatterns() instead of llamasamplerinitgrammar_lazy()
  • Perform warmup runs manually instead of calling deprecated llamasetwarmup() (deprecated in b9704)

See the API reference for complete mappings.