> ## Documentation Index
> Fetch the complete documentation index at: https://docs.portkey.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Thinking Mode

Thinking/Reasoning models represent a new generation of LLMs specifically trained to expose their internal reasoning process. Unlike traditional LLMs that only show final outputs, thinking models like Claude 3.7 Sonnet, OpenAI o1/o3, and Deepseek R1 are designed to "think out loud" - producing a detailed chain of thought before delivering their final response.

These reasoning-optimized models are built to excel in tasks requiring complex analysis, multi-step problem solving, and structured logical thinking. The Prisma AIRS AI Gateway makes these advanced models accessible through a unified API specification that works consistently across providers.

## Supported Thinking Models

The AI Gateway currently supports thinking-enabled models from **Anthropic**, **Google Vertex AI**, **Amazon Bedrock**, **OpenAI**, **Together AI**, and other **OpenAI-compatible providers**.

*Note: If a specific model is not supported for thinking on the AI Gateway, please reach out to us at [support@portkey.ai](mailto:support@portkey.ai).*

## Using Thinking Mode

1. You must set `strict_open_ai_compliance=False` in your headers or client configuration
2. The thinking response is returned in a different format than standard completions
3. For streaming responses, the thinking content is in `response_chunk.choices[0].delta.content_blocks`

<Warning>
  Extended thinking API through the AI Gateway is currently in beta.
</Warning>

### Basic Example

<Tabs>
  <Tab title="OpenAI SDK (JS)">
    ```javascript theme={"system"}
    import OpenAI from 'openai'; // We're using the v4 SDK

    const openai = new OpenAI({
      apiKey: "PORTKEY_API_KEY", // defaults to process.env["OPENAI_API_KEY"],
      baseURL: "https://aigw.portkey.ai/v1",
      defaultHeaders: {
          // defaults to process.env["PORTKEY_API_KEY"]
          "x-portkey-strict-open-ai-compliance": false,
      }
    });

    // Generate a chat completion with thinking mode
    async function getChatCompletionFunctions(){
      const response = await openai.chat.completions.create({
        model: "@anthropic/claude-3-7-sonnet-latest",
        max_tokens: 3000,
        thinking: {
            type: "enabled",
            budget_tokens: 2030
        },
        stream: false,
        messages: [
            {
                role: "user",
                content: [
                    {
                        type: "text",
                        text: "when does the flight from new york to bengaluru land tomorrow, what time, what is its flight number, and what is its baggage belt?"
                    }
                ]
            }
        ],
      });

      console.log(response)
    }
    await getChatCompletionFunctions();
    ```
  </Tab>

  <Tab title="OpenAI SDK (Python)">
    ```python theme={"system"}
    from openai import OpenAI

    openai = OpenAI(
        api_key="PORTKEY_API_KEY",
        base_url="https://aigw.portkey.ai/v1",
        default_headers={"x-portkey-provider": "anthropic", "x-portkey-strict-open-ai-compliance": False}
    )

    response = openai.chat.completions.create(
        model="claude-3-7-sonnet-latest",
        max_tokens=3000,
        thinking={
            "type": "enabled",
            "budget_tokens": 2030
        },
        stream=False,
        messages=[
            {
                "role": "user",
                "content": [
                    {
                        "type": "text",
                        "text": "when does the flight from new york to bengaluru land tomorrow, what time, what is its flight number, and what is its baggage belt?"
                    }
                ]
            }
        ]
    )

    print(response)
    ```
  </Tab>

  <Tab title="cURL">
    ```sh theme={"system"}
    curl "https://aigw.portkey.ai/v1/chat/completions" \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer $PORTKEY_API_KEY" \
      -H "x-api-key: $ANTHROPIC_API_KEY" \
      -H "x-portkey-strict-open-ai-compliance: false" \
      -d '{
        "model": "@anthropic/claude-3-7-sonnet-latest",
        "max_tokens": 3000,
        "thinking": {
          "type": "enabled",
          "budget_tokens": 2030
        },
        "stream": false,
        "messages": [
          {
            "role": "user",
            "content": [
              {
                "type": "text",
                "text": "when does the flight from new york to bengaluru land tomorrow, what time, what is its flight number, and what is its baggage belt?"
              }
            ]
          }
        ]
      }'
    ```
  </Tab>
</Tabs>

## Multi-Turn Conversations

For multi-turn conversations, include the previous thinking content in the conversation history:

<CodeGroup>
  ```js OpenAI NodeJS theme={"system"}
  import OpenAI from 'openai'; // We're using the v4 SDK

  const openai = new OpenAI({
    apiKey: "PORTKEY_API_KEY", // defaults to process.env["OPENAI_API_KEY"],
    baseURL: "https://aigw.portkey.ai/v1",
    defaultHeaders: {
        "x-portkey-provider": "anthropic",
        // defaults to process.env["PORTKEY_API_KEY"]
        "x-portkey-strict-open-ai-compliance": false,
    }
  });

  // Generate a chat completion with streaming
  async function getChatCompletionFunctions(){
    const response = await openai.chat.completions.create({
      model: "@anthropic/claude-3-7-sonnet-latest",
      max_tokens: 3000,
      thinking: {
          type: "enabled",
          budget_tokens: 2030
      },
      stream: false,
      messages: [
          {
              role: "user",
              content: [
                  {
                      type: "text",
                      text: "when does the flight from baroda to bangalore land tomorrow, what time, what is its flight number, and what is its baggage belt?"
                  }
              ]
          },
          {
              role: "assistant",
              content: [
                      {
                          type: "thinking",
                          thinking: "The user is asking several questions about a flight from Baroda (also known as Vadodara) to Bangalore:\n1. When does the flight land tomorrow\n2. What time does it land\n3. What is the flight number\n4. What is the baggage belt number at the arrival airport\n\nTo properly answer these questions, I would need access to airline flight schedules and airport information systems. However, I don't have:\n- Real-time or scheduled flight information\n- Access to airport baggage claim allocation systems\n- Information about specific flights between these cities\n- The ability to look up tomorrow's specific flight schedules\n\nThis question requires current, specific flight information that I don't have access to. Instead of guessing or providing potentially incorrect information, I should explain this limitation and suggest ways the user could find this information.",
                          signature: "EqoBCkgIARABGAIiQBVA7FBNLRtWarDSy9TAjwtOpcTSYHJ+2GYEoaorq3V+d3eapde04bvEfykD/66xZXjJ5yyqogJ8DEkNMotspRsSDKzuUJ9FKhSNt/3PdxoMaFZuH+1z1aLF8OeQIjCrA1+T2lsErrbgrve6eDWeMvP+1sqVqv/JcIn1jOmuzrPi2tNz5M0oqkOO9txJf7QqEPPw6RG3JLO2h7nV1BMN6wE="
                      }
              ]
          },
          {
              role: "user",
              content: "thanks that's good to know, how about to chennai?"
          }
      ],
    });

    console.log(response)

  }
  await getChatCompletionFunctions();
  ```

  ```py OpenAI Python theme={"system"}
  from openai import OpenAI

  openai = OpenAI(
      api_key='PORTKEY_API_KEY',
      base_url="https://aigw.portkey.ai/v1",
      default_headers={"x-portkey-strict-open-ai-compliance": False}
  )

  response = openai.chat.completions.create(
      model="@anthropic/claude-3-7-sonnet-latest",
      max_tokens=3000,
      thinking={
          "type": "enabled",
          "budget_tokens": 2030
      },
      stream=False,
      messages=[
          {
              "role": "user",
              "content": [
                  {
                      "type": "text",
                      "text": "when does the flight from baroda to bangalore land tomorrow, what time, what is its flight number, and what is its baggage belt?"
                  }
              ]
          },
          {
              "role": "assistant",
              "content": [
                      {
                          "type": "thinking",
                          "thinking": "The user is asking several questions about a flight from Baroda (also known as Vadodara) to Bangalore:\n1. When does the flight land tomorrow\n2. What time does it land\n3. What is the flight number\n4. What is the baggage belt number at the arrival airport\n\nTo properly answer these questions, I would need access to airline flight schedules and airport information systems. However, I don't have:\n- Real-time or scheduled flight information\n- Access to airport baggage claim allocation systems\n- Information about specific flights between these cities\n- The ability to look up tomorrow's specific flight schedules\n\nThis question requires current, specific flight information that I don't have access to. Instead of guessing or providing potentially incorrect information, I should explain this limitation and suggest ways the user could find this information.",
                          signature: "EqoBCkgIARABGAIiQBVA7FBNLRtWarDSy9TAjwtOpcTSYHJ+2GYEoaorq3V+d3eapde04bvEfykD/66xZXjJ5yyqogJ8DEkNMotspRsSDKzuUJ9FKhSNt/3PdxoMaFZuH+1z1aLF8OeQIjCrA1+T2lsErrbgrve6eDWeMvP+1sqVqv/JcIn1jOmuzrPi2tNz5M0oqkOO9txJf7QqEPPw6RG3JLO2h7nV1BMN6wE="
                      }
              ]
          },
          {
              "role": "user",
              "content": "thanks that's good to know, how about to chennai?"
          }
      ]
  )

  print(response)
  ```

  ```sh cURL theme={"system"}
  curl "https://aigw.portkey.ai/v1/chat/completions" \
    -H "Content-Type: application/json" \
    -H "Authorization: Bearer $PORTKEY_API_KEY" \
    -H "x-portkey-strict-open-ai-compliance: false" \
    -d '{
      "model": "@anthropic/claude-3-7-sonnet-latest",
      "max_tokens": 3000,
      "thinking": {
        "type": "enabled",
        "budget_tokens": 2030
      },
      "stream": false,
      "messages": [
        {
          "role": "user",
          "content": [
            {
              "type": "text",
              "text": "when does the flight from baroda to bangalore land tomorrow, what time, what is its flight number, and what is its baggage belt?"
            }
          ]
        },
        {
          "role": "assistant",
          "content": [
                  {
                      "type": "thinking",
                      "thinking": "The user is asking several questions about a flight from Baroda (also known as Vadodara) to Bangalore:\n1. When does the flight land tomorrow\n2. What time does it land\n3. What is the flight number\n4. What is the baggage belt number at the arrival airport\n\nTo properly answer these questions, I would need access to airline flight schedules and airport information systems. However, I don't have:\n- Real-time or scheduled flight information\n- Access to airport baggage claim allocation systems\n- Information about specific flights between these cities\n- The ability to look up tomorrow's specific flight schedules\n\nThis question requires current, specific flight information that I don't have access to. Instead of guessing or providing potentially incorrect information, I should explain this limitation and suggest ways the user could find this information.",
                      "signature": "EqoBCkgIARABGAIiQBVA7FBNLRtWarDSy9TAjwtOpcTSYHJ+2GYEoaorq3V+d3eapde04bvEfykD/66xZXjJ5yyqogJ8DEkNMotspRsSDKzuUJ9FKhSNt/3PdxoMaFZuH+1z1aLF8OeQIjCrA1+T2lsErrbgrve6eDWeMvP+1sqVqv/JcIn1jOmuzrPi2tNz5M0oqkOO9txJf7QqEPPw6RG3JLO2h7nV1BMN6wE="
                  }
          ]
        },
        {
          "role": "user",
          "content": "thanks that's good to know, how about to chennai?"
        }
      ]
    }'
  ```
</CodeGroup>

## Understanding Response Format

When using thinking-enabled models, be aware of the special response format:

<Note>
  The assistant's thinking response is returned in the `response_chunk.choices[0].delta.content_blocks` array, not the `response.choices[0].message.content` string.
</Note>

This is especially important for streaming responses, where you'll need to specifically parse and extract the thinking content from the content blocks.

## When to Use Thinking Models

Thinking models are particularly valuable in specific use cases:

* **Complex problem solving:** Break down multi-step tasks such as math, planning, debugging, or technical analysis where the model needs to reason through intermediate steps before answering.
* **High-stakes decision support:** Review contracts, policies, research, or operational runbooks where showing the reasoning path helps users audit the answer and catch missed assumptions.
* **Long-context analysis:** Compare multiple documents, synthesize evidence, or trace dependencies across large inputs where the model benefits from allocating tokens to structured reasoning.
* **Agentic workflows:** Power agents that need to plan tool calls, evaluate options, recover from errors, or explain why they chose a particular next step.
* **Quality-sensitive generation:** Produce answers that need stronger consistency, constraint following, and self-checking, such as code review, architecture recommendations, or data analysis.

For simple classification, extraction, summarization, or low-latency chat flows, a standard chat model is usually more cost-effective. Thinking mode adds reasoning tokens, so tune `budget_tokens` based on the complexity of the task and the amount of reasoning you want returned.

## FAQs

<AccordionGroup>
  <Accordion title="Can I use thinking mode with any model?">
    No, thinking mode is only available on specific reasoning-optimized models. Currently, this includes Claude 3.7 Sonnet and will expand to other models as they become available.
  </Accordion>

  <Accordion title="Does thinking mode increase token usage?">
    Yes, enabling thinking mode will increase your token usage since the model is generating additional content for its reasoning process. The `budget_tokens` parameter lets you control the maximum tokens allocated to thinking.
  </Accordion>

  <Accordion title="Do I need to handle the response differently for thinking mode?">
    Yes, particularly for streaming responses. The thinking content is returned in the `content_blocks` array rather than the standard content field, so you'll need to adapt your response parsing logic.
  </Accordion>

  <Accordion title="Why do I need to set strict_open_ai_compliance to false?">
    The thinking mode response format extends beyond the standard OpenAI completion schema. Setting `strict_open_ai_compliance` to false allows the AI Gateway to return this extended format with the thinking content.
  </Accordion>
</AccordionGroup>


## Related topics

- [Together AI](/docs/aigw/integrations/llms/together-ai.md)
- [Messages](/docs/aigw/product/ai-gateway/messages-api.md)
- [Enterprise Gateway](/docs/aigw/changelog/enterprise.md)
- [Google Vertex AI](/docs/aigw/integrations/llms/vertex-ai.md)
- [AWS Bedrock](/docs/aigw/integrations/llms/bedrock/aws-bedrock.md)
