All models
MetaFastAvailable

Llama 3.3 70B

meta-llama/llama-3.3-70b

Reliable workhorse for chat and RAG

fasttools
Input
$0.11
per 1M tokens
Output
$0.34
per 1M tokens
Typical reply
≈ $0.00028
1k tokens in, 500 out
Context
128K
Context
Speed
120 tok/s
*
Time to first token
0.5 s
*

* Speed and latency are approximate.

How to use

Use the model ID in the model field of any OpenAI-compatible request.

curl https://zelvimou.com/v1/chat/completions \
  -H "Authorization: Bearer sk-zelvimo-YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/llama-3.3-70b",
    "messages": [{"role": "user", "content": "Hello! What can you do?"}]
  }'

Replace sk-zelvimo-YOUR_KEY with your key from the API keys page.

Similar models

  • Alibaba · Fast

    Qwen3 Coder

    Open model for agents in your IDE

    $0.32/ $1.05

    Input / 1M · Output / 1M

    Context 256K140 tok/s
  • Anthropic · Fast

    Claude Haiku 4.5

    Temporarily unavailable

    Fast Claude for high-volume tasks

    $1.00/ $5.00

    Input / 1M · Output / 1M

    Context 200K150 tok/s