zainhas/skills · Archived

together-batch-inference

High-volume, asynchronous offline inference at up to 50% lower cost via Together AI's Batch API.

First seen Mar 30, 2026

Installation

$ npx skills add zainhas/skills --skill together-batch-inference

Summary

  • High-volume, asynchronous offline inference at up to 50% lower cost via Together AI's Batch API.
  • Prepare JSONL inputs, upload files, create jobs, poll status, and download outputs.
  • Reach for it whenever the user needs non-interactive bulk inference rather than real-time chat or evaluation jobs.

Stronger alternatives

This repository is archived — consider an actively maintained alternative.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from zainhas/skills · top by installs.

npx skills add zainhas/skills

Browse all from zainhas/skills

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 1
License LICENSE
Default branch main
Status Archived

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 3,290 B
  • docs SUMMARY.md 327 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 1 installs

SKILL.md

Together Batch Inference

Overview

Use Together AI's Batch API for large offline workloads where latency is not the primary concern.

Typical fits:

  • bulk classification
  • synthetic data generation
  • dataset transformations
  • large summarization or enrichment jobs
  • low-cost asynchronous inference

When This Skill Wins

  • The user has many independent requests to run
  • A JSONL request file is acceptable
  • Turnaround time can be minutes or hours instead of seconds
  • Lower cost matters more than immediate interactivity

Hand Off To Another Skill

  • Use together-chat-completions for real-time requests or tool-calling apps
  • Use together-evaluations for managed LLM-as-a-judge workflows
  • Use together-embeddings for retrieval-specific vector generation

Quick Routing

  • End-to-end batch workflow

- Start with [scripts/batchworkflow.py](scripts/batchworkflow.py) or [scripts/batchworkflow.ts](scripts/batchworkflow.ts)

  • Request format, status model, and result downloads

- Read [references/api-reference.md](references/api-reference.md)

  • Operational guidance and batch sizing

- Read [references/api-reference.md](references/api-reference.md)

Workflow

  1. Build a JSONL file where each line contains custom_id and body.
  2. Upload the file with purpose="batch-api".
  3. Create the batch with inputfileid=... and the target endpoint.
  4. Poll until the job is terminal.
  5. Download output and error files, then reconcile by custom_id.

High-Signal Rules

  • Python scripts require the Together v2 SDK (together>=2.0.0). If the user is on an older version, they must upgrade first: uv pip install --upgrade "together>=2.0.0".
  • Use inputfileid, not legacy file parameters.
  • Keep custom_id stable and meaningful so result reconciliation is easy.
  • Batch is for independent requests. If the workload depends on shared conversation state, it is probably the wrong tool.
  • Always inspect the error file in addition to the success output.
  • client.batches.create() returns a wrapper; access the batch object via response.job (e.g., response.job.id). client.batches.retrieve() returns the batch object directly.
  • For classification or labeling workloads, set max_tokens low (e.g., 4), use temperature: 0, and constrain the system prompt to return only the label. This minimizes output tokens and cost.
  • Small batches (under 1K requests) typically complete in minutes. The 24-hour completion window is a maximum, not typical.

Resource Map

  • API reference and operational guidance: [references/api-reference.md](references/api-reference.md)
  • Python workflow: [scripts/batchworkflow.py](scripts/batchworkflow.py)
  • TypeScript workflow: [scripts/batchworkflow.ts](scripts/batchworkflow.ts)

Official Docs