> ## Documentation Index
> Fetch the complete documentation index at: https://concentrate.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# API Introduction

> Unified API for accessing 150+ AI models through a single interface

## Welcome to [Concentrate AI](https://concentrate.ai/?utm_source=dedo)

The Concentrate AI Responses API provides a unified interface for interacting with multiple AI model providers. Access GPT 5.6, Claude Opus 5, Claude Fable 5, Gemini 3.1 Pro, Grok 4.20, and many other models through a single, normalized API with automatic routing and credit tracking.

<CardGroup cols={2}>
  <Card title="Quickstart" icon="rocket" href="/docs/getting-started/quickstart">
    Get started with your first API request in minutes
  </Card>

  <Card title="API Reference" icon="code" href="/docs/api-reference/endpoint/create-response">
    View detailed endpoint documentation
  </Card>

  <Card title="Claude Code Setup" icon="terminal" href="/docs/integrations/claude-code">
    Use Claude Code with any model in our [model fortress](https://concentrate.ai/models)
  </Card>

  <Card title="Cursor Setup" icon="pen-to-square" href="/docs/integrations/cursor">
    Use models on Concentrate in Cursor
  </Card>
</CardGroup>

## Key Features

<AccordionGroup>
  <Accordion title="Unified Interface" icon="layer-group">
    One API format works across all providers. No need to learn different request/response formats for OpenAI, Anthropic, Google, or other providers.
  </Accordion>

  <Accordion title="Automatic Routing" icon="route">
    Use `model: "auto"` to automatically select the best model based on cost, performance, or latency. The API intelligently routes your requests based on real-time metrics.
  </Accordion>

  <Accordion title="Multi-Provider Support" icon="server">
    Access models from OpenAI, Anthropic, Google Vertex, Google AI Studio, AWS Bedrock, Azure, Azure AI Foundry, xAI, Cohere, Mistral, Alibaba Cloud, Cloudflare, z.ai, MiniMax, DeepSeek, DeepInfra, Novita, Blue Lobster, and Concentrate through a single endpoint.
  </Accordion>

  <Accordion title="Streaming Responses" icon="water">
    Enable real-time streaming via Server-Sent Events (SSE) for a responsive user experience. Works consistently across all providers.
  </Accordion>

  <Accordion title="Credit Tracking" icon="credit-card">
    Built-in usage tracking and billing integration. Monitor token usage, costs, and set spending limits.
  </Accordion>
</AccordionGroup>

## Authentication

All API requests require authentication using an API key. Get your API key from the [Concentrate.ai dashboard](https://concentrate.ai).

Include your API key in the `Authorization` header:

```bash theme={null}
Authorization: Bearer YOUR_API_KEY
```

<Warning>
  Keep your API key secure. Never share it publicly or commit it to version control.
</Warning>

## Base URL

```text theme={null}
https://api.concentrate.ai/v1
```

## Supported Models

The API supports **157+ models** from **21 authors** across **19 providers**:

<Tabs>
  <Tab title="OpenAI">
    **26 models**

    * GPT 5.6 Sol, GPT 5.6 Terra, GPT 5.6 Luna
    * GPT 5.5, GPT 5.4, GPT 5.4 Pro, GPT 5.4 Mini, GPT 5.4 Nano
    * GPT 5.3 Codex, GPT 5.2 Codex, GPT 5.1 Codex Max, GPT 5.1 Codex Mini
    * GPT 5.2, GPT 5.1, GPT 5, GPT 5 Mini, GPT 5 Nano
    * GPT 4.1, GPT 4.1 Mini, GPT 4o, GPT 4o Mini
    * o1 (reasoning)
    * gpt-oss 120B, gpt-oss 20B, gpt-oss Safeguard 120B / 20B
  </Tab>

  <Tab title="Anthropic">
    **12 models**

    * Claude Fable 5, Claude Sonnet 5
    * Claude Opus 5, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Opus 4.5, Claude Opus 4.1
    * Claude Sonnet 4.6, Claude Sonnet 4.5, Claude Sonnet 4
    * Claude Haiku 4.5
  </Tab>

  <Tab title="Google">
    **12 models**

    * Gemini 3.1 Pro Preview, Gemini 3.1 Flash Lite Preview
    * Gemini 3.5 Flash, Gemini 3 Flash Preview
    * Gemini 2.5 Pro, Gemini 2.5 Flash, Gemini 2.5 Flash Lite
    * Gemma 4 31B, Gemma 4 26B, Gemma 3 27B / 12B / 4B
  </Tab>

  <Tab title="xAI">
    **14 models**

    * Grok 4.20 Multi-Agent, Grok 4.20 Reasoning, Grok 4.20 Non Reasoning
    * Grok 4.1 Fast Reasoning, Grok 4.1 Fast
    * Grok 4.5, Grok 4.3, Grok 4, Grok 4 (0709), Grok 4 Fast Reasoning, Grok 4 Fast
    * Grok 3, Grok 3 Mini
    * Grok Build 0.1
  </Tab>

  <Tab title="Meta">
    **11 models**

    * Llama 4 Maverick, Llama 4 Scout
    * Llama 3.3 70B, Llama 3.2 (1B / 3B / 11B / 90B), Llama 3.1 (8B / 70B)
    * Llama 3 70B, Llama 3 8B
  </Tab>

  <Tab title="Mistral">
    **13 models**

    * Mistral Large 3, Mistral Medium 3.1, Mistral Medium 3
    * Magistral Medium 1.2, Magistral Small 1.2 (reasoning)
    * Codestral, Devstral 2
    * Mistral Small 3.2, Mistral Small 3.1, Mistral Nemo
    * Ministral 3 (3B / 8B / 14B)
  </Tab>

  <Tab title="Alibaba">
    **24 models**

    * Qwen3.7 Max, Qwen3.7 Plus
    * Qwen3.6 Plus, Qwen3.6 Flash, Qwen3.6 35B
    * Qwen3.5 Plus, Qwen3.5 Flash, Qwen3.5 397B / 122B / 35B
    * Qwen3 Max, Qwen3 Coder Plus / Flash / Next / 30B
    * Qwen3 30B / 32B, Qwen3 Next 80B, Qwen3 VL 235B / Plus / Flash
    * QwQ 32B, Qwen Plus, Qwen Flash
  </Tab>

  <Tab title="Others">
    * **DeepSeek** (7): DeepSeek V4 Pro, V4 Flash, V3.2, V3.1, R1, R1 0528, R1 Distill 32B
    * **z.ai** (9): GLM-5, GLM-5.2, GLM-5.1, GLM-4.7, GLM-4.7 Flash, GLM-4.6, GLM-4.6v, GLM-4.5, GLM-4.5v
    * **MiniMax** (8): MiniMax M3, M2.7, M2.5, M2.1, M2, plus Highspeed variants
    * **Moonshot AI** (4): Kimi K2.7 Code, K2.6, K2.5, K2 Thinking
    * **Amazon** (5): Nova Premier, Nova Pro, Nova Lite, Nova Micro, Nova 2 Lite
    * **Cohere** (2): Command A, Command A Vision
    * **NVIDIA** (2): Nemotron 3 120B, Nemotron 3 Nano Omni
    * **Writer** (2): Palmyra X5, Palmyra X4
    * **AI21 Labs** (2): Jamba 1.5 Large, Jamba 1.5 Mini
    * **Xiaomi**: MiMo V2.5
    * **StepFun AI**: Step 3.5 Flash
    * **IBM**: Granite Micro
    * **Tencent**: Hunyuan Hy3
    * **Concentrate**: Redact v1
  </Tab>
</Tabs>

<Info>
  Check the [**Model Fortress page**](https://concentrate.ai/models) for complete listings, live pricing, and per-provider availability. You can also query [`/v1/models`](/docs/api-reference/endpoint/list-models) for real-time data.
</Info>

## Model Selection

You can specify models in three ways:

1. **Model name only**: `"gpt-5.6-sol"` - Automatic provider routing
2. **Provider prefix**: `"openai/gpt-5.6-sol"` - Specific provider
3. **Auto routing**: `"auto"` - Let the API choose based on your criteria

<CodeGroup>
  ```json Model Name theme={null}
  {
    "model": "gpt-5.6-sol",
    "input": "Hello, world!"
  }
  ```

  ```json Provider Prefix theme={null}
  {
    "model": "anthropic/claude-opus-4-8",
    "input": "Hello, world!"
  }
  ```

  ```json Auto Routing theme={null}
  {
    "model": "auto",
    "input": "Hello, world!",
    "routing": {
      "strategy": "min",
      "metric": "cost"
    }
  }
  ```
</CodeGroup>

## Quick Example

<CodeGroup>
  ```bash cURL theme={null}
  curl https://api.concentrate.ai/v1/responses \
    -H "Content-Type: application/json" \
    -H "Authorization: Bearer YOUR_API_KEY" \
    -d '{
      "model": "gpt-5.6-sol",
      "input": "What is the capital of France?"
    }'
  ```

  ```python Python theme={null}
  import requests

  response = requests.post(
      "https://api.concentrate.ai/v1/responses",
      headers={
          "Authorization": "Bearer YOUR_API_KEY",
          "Content-Type": "application/json"
      },
      json={
          "model": "gpt-5.4",
          "input": "What is the capital of France?"
      }
  )

  print(response.json())
  ```

  ```javascript JavaScript theme={null}
  const response = await fetch("https://api.concentrate.ai/v1/responses", {
    method: "POST",
    headers: {
      "Authorization": "Bearer YOUR_API_KEY",
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      model: "gpt-5.4",
      input: "What is the capital of France?"
    })
  });

  const data = await response.json();
  console.log(data);
  ```

  ```typescript TypeScript theme={null}
  interface ResponseRequest {
    model: string;
    input: string | Array<{role: string; content: string}>;
    stream?: boolean;
    temperature?: number;
    max_output_tokens?: number;
  }

  const response = await fetch("https://api.concentrate.ai/v1/responses", {
    method: "POST",
    headers: {
      "Authorization": "Bearer YOUR_API_KEY",
      "Content-Type": "application/json"
    },
    body: JSON.stringify({
      model: "gpt-5.4",
      input: "What is the capital of France?"
    } as ResponseRequest)
  });

  const data = await response.json();
  console.log(data);
  ```
</CodeGroup>

## Response Format

All responses follow a normalized format regardless of provider:

```json theme={null}
{
  "id": "resp_abc123",
  "created_at": 1702934400,
  "status": "completed",
  "model": "openai/gpt-5.6-sol",
  "output": [
    {
      "type": "message",
      "id": "msg_xyz789",
      "status": "completed",
      "role": "assistant",
      "content": [
        {
          "type": "output_text",
          "text": "The capital of France is Paris."
        }
      ]
    }
  ],
  "usage": {
    "input_tokens": 12,
    "input_tokens_details": {
      "cached_tokens": 0
    },
    "output_tokens": 8,
    "output_tokens_details": {
      "reasoning_tokens": 0
    },
    "total_tokens": 20
  }
}
```

## Error Handling

The API uses standard HTTP status codes:

| Status Code | Description                             |
| ----------- | --------------------------------------- |
| `200`       | Successful request                      |
| `400`       | Bad request - Invalid parameters        |
| `401`       | Unauthorized - Invalid API key          |
| `402`       | Payment required - Insufficient credits |
| `424`       | Failed dependency - Provider error      |
| `500`       | Internal server error                   |

<Card title="View Error Examples" icon="triangle-exclamation" href="/docs/api-reference/endpoint/errors">
  See detailed error response formats and troubleshooting
</Card>

## Rate Limits

Rate limits are applied per API key and are based on your subscription tier. Limits are enforced using token bucket algorithm with per-minute windows.

<Info>
  Contact support to increase your rate limits or discuss enterprise pricing.
</Info>

## Next Steps

<CardGroup cols={2}>
  <Card title="Quickstart Guide" icon="play" href="/docs/getting-started/quickstart">
    Make your first API call
  </Card>

  <Card title="Create Response" icon="message" href="/docs/api-reference/endpoint/create-response">
    Full endpoint documentation
  </Card>

  <Card title="Streaming" icon="water" href="/docs/api-reference/endpoint/streaming">
    Learn about streaming responses
  </Card>

  <Card title="Auto Routing" icon="route" href="/docs/api-reference/endpoint/auto-routing">
    Automatic model selection
  </Card>
</CardGroup>
