> ## Documentation Index
> Fetch the complete documentation index at: https://docs.portkey.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Multimodal Capabilities

<Info>
  This feature is available on all Prisma AIRS AI Gateway plans.
</Info>

The Gateway is your unified interface for **multimodal models**, along with chat, text, and embedding models.

Using the Gateway, you can call `vision`, `audio (text-to-speech & speech-to-text)`, `image generation` and other multimodal models from multiple providers (like `OpenAI`, `Anthropic`, `Stability AI`, etc.) — all using the familiar OpenAI signature.

#### Explore the AI Gateway's Multimodal capabilities below:

<Card title="Vision" href="/docs/aigw/product/ai-gateway/multimodal-capabilities/vision" />

<Card title="Image Generation" href="/docs/aigw/product/ai-gateway/multimodal-capabilities/image-generation" />

<Card title="Function Calling" href="/docs/aigw/product/ai-gateway/multimodal-capabilities/function-calling" />

<Card title="Speech-to-Text" href="/docs/aigw/product/ai-gateway/multimodal-capabilities/speech-to-text" />

<Card title="Text-to-Speech" href="/docs/aigw/product/ai-gateway/multimodal-capabilities/text-to-speech" />


## Related topics

- [Universal API](/docs/aigw/product/ai-gateway/universal-api.md)
- [Google Gemini](/docs/aigw/integrations/llms/gemini.md)
- [Features](/docs/aigw/introduction/feature-overview.md)
- [Supported Endpoints & Capabilities](/docs/aigw/product/guardrails/capabilities.md)
- [Image Generation](/docs/aigw/product/ai-gateway/multimodal-capabilities/image-generation.md)
