Mistral Large 4 went live on the Mistral API on October 6, 2026, three weeks before its open weights. If you want to try the 1-trillion-parameter “Le Chonk” now, the API is the only way in, and right now it is also the cheapest way: Mistral lists it at $0.68 per million input tokens and $2.09 per million output tokens during the public preview, half the $1.36 / $4.18 list price.
This guide gets you from zero to a working first call in about five minutes, then covers the parts that trip people up: reasoning chunks, image input, function calling, JSON output, and cost. Every request can be saved and replayed in Apidog so you can compare Large 4 with whatever model you run today.
New to the model itself? Read Mistral Is Back: Le Chonk Beats GPT-6 Astra and Claude at Cyber first for the benchmarks and the catch behind the cyber headline.
What you need
| Item | Value |
|---|---|
| Base URL | https://api.mistral.ai/v1 |
| Auth | Authorization: Bearer $MISTRAL_API_KEY |
| Model ID |
mistral-large-4 (alias mistral-large-4-0) |
| Main endpoint | POST /v1/chat/completions |
| Context window | 1M tokens |
| Input types | Text, images |
| Python SDK | pip install mistralai |
| TypeScript SDK | npm install @mistralai/mistralai |
Step 1: Get an API key
- Sign in to Mistral Studio (formerly La Plateforme).
- Open API Keys and create a new key. Name it for its environment, such as
local-devorci-staging. - Copy the key immediately. Studio will not show it again.
- Export it in your shell:
export MISTRAL_API_KEY="your-key-here"
Keep the key out of source control. If you are wiring it into several tools, see API key management best practices for rotation and scoping guidance.
Step 2: Make your first call
Start with a plain curl request:
curl https://api.mistral.ai/v1/chat/completions \
-H "Authorization: Bearer $MISTRAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mistral-large-4",
"messages": [
{
"role": "user",
"content": "Give me three edge cases to test on a pagination API."
}
]
}'
A successful response includes the answer in choices[0].message.content and a usage object containing prompt_tokens, completion_tokens, and total_tokens.
Common first-call failures:
- 401: The API key is missing, incorrect, or was not exported in the current shell.
- 404 for the model: Check the model ID for typos.
The same call in Python
Install the SDK:
pip install mistralai
Then call the API:
import os
from mistralai import Mistral
client = Mistral(api_key=os.environ["MISTRAL_API_KEY"])
response = client.chat.complete(
model="mistral-large-4",
messages=[
{
"role": "user",
"content": "Give me three edge cases to test on a pagination API.",
}
],
)
print(response.choices[0].message.content)
And in TypeScript
Install the SDK:
npm install @mistralai/mistralai
Then make the request:
import { Mistral } from "@mistralai/mistralai";
const client = new Mistral({ apiKey: process.env.MISTRAL_API_KEY });
const response = await client.chat.complete({
model: "mistral-large-4",
messages: [
{
role: "user",
content: "Give me three edge cases to test on a pagination API.",
},
],
});
console.log(response.choices[0].message.content);
Step 3: Save it in Apidog
Typing curl commands gets old as soon as you begin comparing models. Save the request in Apidog:
- Create a new HTTP request:
POST https://api.mistral.ai/v1/chat/completions. - Add an environment variable named
MISTRAL_API_KEY. - Set the authorization header:
Authorization: Bearer {{MISTRAL_API_KEY}}
- Paste the JSON body from Step 2 and select Send.
- Duplicate the request, change
modelto the model you use today—for example,mistral-medium-3-5—and run both requests.
You now have two saved requests using the same prompt. Apidog shows the response body, status, timing, and size for each request, making it easier to compare answer quality, latency, and usage token counts without writing a comparison script.
Add a post-response assertion that choices[0].message.content is not empty. That gives you a smoke test you can rerun whenever Mistral updates the preview.
Step 4: Turn reasoning on and off
Large 4 is a hybrid model: the same model handles fast answers and step-by-step reasoning. Control this behavior with reasoning_effort.
| Value | Behavior | Use it for |
|---|---|---|
"none" |
Minimal thinking, with no thinking chunk in the response | Chat, extraction, classification, and latency-sensitive requests |
"high" |
Full thinking chunk before the final answer | Debugging, multi-step planning, math, and code review |
Example with high reasoning effort:
curl https://api.mistral.ai/v1/chat/completions \
-H "Authorization: Bearer $MISTRAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mistral-large-4",
"messages": [
{
"role": "user",
"content": "Our API returns 200 with an empty body under load. List likely causes in order of probability."
}
],
"reasoning_effort": "high"
}'
This is the part that breaks parsers. With reasoning_effort: "high", message.content is no longer a string. It becomes a list of chunks:
- A
thinkingchunk containing the reasoning trace. - A
textchunk containing the final answer.
Extract the text chunks explicitly:
response = client.chat.complete(
model="mistral-large-4",
messages=[
{
"role": "user",
"content": "Why would a 200 response have an empty body?",
}
],
reasoning_effort="high",
)
content = response.choices[0].message.content
if isinstance(content, str):
answer = content
else:
answer = "".join(chunk.text for chunk in content if chunk.type == "text")
print(answer)
Thinking tokens are billed as output tokens, so "high" costs more per request. Default to "none" and use "high" only where the extra reasoning is useful.
Step 5: Send an image
Large 4 is natively multimodal, with a 1.6B-parameter vision encoder. Send images as content parts alongside text:
curl https://api.mistral.ai/v1/chat/completions \
-H "Authorization: Bearer $MISTRAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mistral-large-4",
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "This is a screenshot of our API error dashboard. Which endpoint is failing most and what is the error code?"
},
{
"type": "image_url",
"image_url": "https://example.com/dashboard.png"
}
]
}
]
}'
For local files, use a base64 data URL:
{
"type": "image_url",
"image_url": "data:image/png;base64,<encoded>"
}
Mistral reports that Large 4 scores 42% on the Dense 200 visual-grounding benchmark, just ahead of GPT-6 Astra’s 41%, so screenshots of dashboards, charts, and UI states are a reasonable fit.
Step 6: Function calling
Function calling is where Large 4’s agent benchmarks—59.9% on AutomationBench—become useful. You provide tool definitions, the model decides whether to invoke one, and your application runs the actual operation.
tools = [
{
"type": "function",
"function": {
"name": "get_order_status",
"description": "Look up the status of an order by its ID.",
"parameters": {
"type": "object",
"properties": {
"order_id": {
"type": "string",
"description": "The order ID, e.g. ORD-1042",
}
},
"required": ["order_id"],
},
},
}
]
messages = [{"role": "user", "content": "Where is order ORD-1042?"}]
response = client.chat.complete(
model="mistral-large-4",
messages=messages,
tools=tools,
tool_choice="auto",
)
tool_call = response.choices[0].message.tool_calls[0]
print(tool_call.function.name, tool_call.function.arguments)
Run the requested function in your application, then send its result back with the matching tool_call_id:
import json
result = {
"order_id": "ORD-1042",
"status": "shipped",
"eta": "2026-10-09",
}
messages.append(response.choices[0].message)
messages.append({
"role": "tool",
"name": "get_order_status",
"content": json.dumps(result),
"tool_call_id": tool_call.id,
})
final = client.chat.complete(
model="mistral-large-4",
messages=messages,
tools=tools,
)
print(final.choices[0].message.content)
The tool schema uses plain JSON Schema. If your API already has an OpenAPI specification, reuse each operation’s request schema in parameters. Designing and maintaining the spec in Apidog can help keep your tool definitions aligned with the real API.
Step 7: Get JSON back
For machine-readable output, set response_format:
curl https://api.mistral.ai/v1/chat/completions \
-H "Authorization: Bearer $MISTRAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mistral-large-4",
"messages": [
{
"role": "user",
"content": "Extract method, path and status code from: GET /v1/users/42 returned 404. Reply in JSON."
}
],
"response_format": {
"type": "json_object"
}
}'
Mention JSON in both the prompt and response_format.
For strict output shapes, Mistral also supports:
{
"type": "json_schema",
"json_schema": {}
}
Provide the full schema in json_schema. In Apidog, add a JSON Schema assertion on the response so an output-shape change fails loudly instead of breaking a downstream service.
What it costs
| Usage | Preview price | List price |
|---|---|---|
| Input, per 1M tokens | $0.68 | $1.36 |
| Cached input, per 1M tokens | $0.07 | $0.14 |
| Output, per 1M tokens | $2.09 | $4.18 |
Consider an agent making 10,000 calls per day. Each call uses 3,000 input tokens—mostly a cached system prompt and tools—and 500 output tokens.
- Input: 30M tokens. If 2,500 of each 3,000 input tokens are cached, that is 25M cached tokens at $0.07 and 5M fresh tokens at $0.68: about $5.15/day.
- Output: 5M tokens at $2.09: about $10.45/day.
- Total: approximately $15.60/day at preview pricing, or about $31 at list price.
The same workload on GPT-6 Astra at $10 / $50 per million tokens, before caching discounts, would cost several hundred dollars per day. Mistral has not said when preview pricing ends, so budget against the list price.
Common errors
| Error | Likely cause | Fix |
|---|---|---|
401 Unauthorized |
Missing or incorrect key | Check echo $MISTRAL_API_KEY and confirm the Bearer prefix |
404 / invalid model |
Typo in the model ID | Use mistral-large-4 exactly |
422 Unprocessable Entity |
Malformed request body, often an invalid tools schema |
Validate the JSON Schema in each tool’s parameters
|
429 Too Many Requests |
Rate limit for your workspace tier | Back off and retry, or raise limits in Studio |
| Answer prints as a list |
reasoning_effort: "high" returns chunks |
Extract the text chunk as shown in Step 4 |
FAQ
Is Mistral Large 4 OpenAI-compatible?
The request shape is very close: model, messages, tools, tool_choice, and response_format work as expected. Use the Mistral SDKs or plain HTTP to be safe. Reasoning output uses Mistral’s own chunk format.
When can I run it locally?
Mistral says the weights ship by the end of October 2026. At 1.05T total parameters, it needs multi-GPU server hardware. The run Mistral 3 locally guide covers tooling for smaller models in the meantime.
Is the preview stable enough for production?
Not yet. The model is labeled public preview and may change before the weights release. Pin your tests, rerun them when Mistral updates the model, and keep a fallback model configured.
Can I use Large 4 with my existing Mistral code?
Yes. The base URL, authentication, and SDK remain the same. Change the model string to mistral-large-4. If you are coming from Medium 3.5, see the Mistral Medium 3.5 API guide for the parts that carry over.
Wrap-up
Five minutes gets you a working call. Spend the next hour running real prompts against Large 4 and your current model side by side. Save both requests in Apidog, add assertions for status and response shape, and you can quickly determine whether Le Chonk earns a place in your stack while preview pricing is still half off.
Top comments (0)