Skip to main content
The AI Gateway provides a robust and secure gateway to integrate various Large Language Models (LLMs) into applications, including Deepinfra’s hosted models. With the AI Gateway, take advantage of features like fast AI gateway access, observability, prompt management, and more, while securely managing API keys through Model Catalog.

Quick Start

Get Deepinfra working in 3 steps:
Tip: You can also send x-portkey-provider: @deepinfra as a header and use just model="nvidia/Nemotron-4-340B-Instruct" in the request.

Add Provider in Model Catalog

  1. Go to Model Catalog → Add Provider
  2. Select Deepinfra
  3. Choose existing credentials or create new by entering your Deepinfra API key
  4. Name your provider (e.g., deepinfra-prod)

Complete Setup Guide →

See all setup options, code examples, and detailed instructions

Supported Endpoints

Tool Calling

DeepInfra supports tool calling (function calling) for compatible models. Use the standard OpenAI tools format:

Supported Models

Deepinfra hosts a wide range of open-source models for text generation. View the complete list:

Deepinfra Models

Browse all available models on Deepinfra
Popular models include:
  • nvidia/Nemotron-4-340B-Instruct
  • meta-llama/Meta-Llama-3.1-405B-Instruct
  • Qwen/Qwen2.5-72B-Instruct

Next Steps

Add Metadata

Add metadata to your Deepinfra requests

Gateway Configs

Add gateway configs to your Deepinfra requests

Tracing

Trace your Deepinfra requests

Fallbacks

Setup fallback from OpenAI to Deepinfra
Last modified on September 15, 2026