Skip to main content
AI Gateway lets you send a single request that fan‑outs to hundreds—or millions—of completions. Choose the mode that best fits cost, latency, and provider support.

Choose Your Batching Mode

Quick rule of thumb →
Need low latency or multi‑provider batching? AI Gateway Batch API. Otherwise, stick with the provider’s native batch for cost savings.

Before You Start

Have the following ready to start making batch requests:
  1. Strata Cloud Manager account & API key.
  2. Data Service to be enabled — required for AI Gateway Managed Batching or when cost-attribution is needed.
  3. Provider credentials for each downstream model (OpenAI key, Bedrock IAM role, etc.).
  4. A AI Gateway File (input_file_id) - required only when using the AI Gateway Batch API (Mode #2). See Files to upload one.
  5. Optional: Familiarity with the Create Batch OpenAPI spec.

Provider Batch API Mode

Used to run batch jobs with the provider’s native batch endpoint. Providers usually offer a cheaper rate for batch jobs, but you’ll be limited by the provider’s quota and limits. Most completion windows are about 24 hours.
Polling for batch status: The AI Gateway is stateless and does not poll for completion status of batches on the provider side. You must poll the batch status manually using the unified API with the same signature for all supported providers. See Retrieve Batch for details.

Quickstart (OpenAI example)

🔗 Full schema: see the OpenAPI reference.

Supported Providers & Endpoints

Need to redact PII or filter requests/responses on a provider batch? See Guardrails for Batches — pass a config with input_guardrails/output_guardrails via portkey_options and the AI Gateway applies them around the upstream batch.

Defaults & Limits


AI Gateway Managed Batching

AI Gateway Managed Batching is a feature that allows you to manage batches across multiple providers with minimal effort and a unified API. Read more about AI Gateway files here.

How It Works

  1. Submit a batch request to the AI Gateway with the gateway file and provider information.
    • Batch requests respect metadata, budgets, and other batch request parameters.
  2. The AI Gateway will automatically upload and start the batch with provider.
  3. The AI Gateway will periodically check the batch status and update the batch status in the AI Gateway.
  4. Once the batch is completed, the AI Gateway reads the batch output for the following details, which are added to your analytics.
    • Token count
    • Cost
    • Success Request count
    • Failed Request count
    • Total Request count
Note: The automatic status polling and analytics described above apply only to AI Gateway Managed Batching. For Provider Batch API Mode (Unified Batch Inference), you must poll the batch status manually.

AI Gateway Custom Batching ⭐️

The AI Gateway custom batching is a feature that allows you to batch requests to the provider which doesn’t have a native batch endpoint.
The AI Gateway custom batching is not a discounted rate.

How It Works

Set completion_window to immediate and the AI Gateway aggregates your requests in memory, then fires them to the target provider in fixed buckets. Coming soon: configurable batch_size, batch_interval, and max_retries.

Quickstart (provider‑agnostic)

Because the AI Gateway orchestrates the batching, this works even for providers without a native batch endpoint.

Response & Monitoring

Identical to Provider mode; the difference is that provider_job_id is absent and cost is computed from individual calls.

About AI Gateway Files

AI Gateway Files are files uploaded to the gateway that are then automatically uploaded to the provider. They’re useful when you want to make multiple batch completions using the same file. The AI Gateway will:
  • Automatically upload the file to the provider on your behalf
  • Reuse the content in your batch requests
  • Check batch progress and provide post-batch analysis including token and cost calculations
  • Make batch outputs available via the GET /batches/<batch_id>/output endpoint

Error Handling & Retries


Security & IAM

  • Files are encrypted at rest (AES‑256) and custom encryption key is supported, if required.
  • The AI Gateway uploads on your behalf using least‑privilege scoped credentials; no long‑lived secrets are stored.
  • Access to batch status & outputs is gated by your workspace role (completions.write).

Glossary


Roadmap

  • Custom batch_size, batch_interval, max_retries (Q3 2025)
  • Real‑time progress webhooks
  • UI for canceling or pausing jobs
Last modified on September 15, 2026