Skip to main content

AI Inference

The AI Inference resource is the OpenAI-compatible gateway for chat completions, text completions, embeddings, and audio. Reference a model by its vendor id; the gateway resolves the provider key and routes the call.

Access it as client.ai.inference on a Strongly client, or the same path on AsyncStrongly with await. All methods exist on both with identical signatures.

Methods​

All methods​

cancel_generation​

cancel_generation(*, job_id: str) -> GenerationJob

Cancel an async generation job.

chat_completion​

chat_completion(*, model: str, messages: Sequence[dict[str, Any] | ChatMessage], stream: bool = False, max_tokens: int | None = None, temperature: float = 0.7, top_p: float = 1.0, stop: str | list[str] | None = None, **kwargs) -> ChatCompletion | Iterator[StreamChunk]

Send a chat completion request.

Parameters

  • model (str): The model ID to use.
  • messages (Sequence[dict[str, Any] | ChatMessage]): List of chat messages.
  • stream (bool, optional): If True, return an iterator of StreamChunk objects.
  • max_tokens (int | None, optional): Maximum tokens to generate.
  • temperature (float, optional): Sampling temperature.
  • top_p (float, optional): Nucleus sampling parameter.
  • stop (str | list[str] | None, optional): Stop sequence(s).
  • **kwargs (Any): Additional parameters forwarded to the API.

Returns

  • ChatCompletion | Iterator[StreamChunk]: ChatCompletion if stream=False, Iterator[StreamChunk] if stream=True.

completion​

completion(*, model: str, prompt: str, stream: bool = False, max_tokens: int | None = None, temperature: float = 0.7, **kwargs) -> Completion | Iterator[StreamChunk]

Send a text completion request.

Parameters

  • model (str): The model ID to use.
  • prompt (str): The text prompt.
  • stream (bool, optional): If True, return an iterator of StreamChunk objects.
  • max_tokens (int | None, optional): Maximum tokens to generate.
  • temperature (float, optional): Sampling temperature.
  • **kwargs (Any): Additional parameters forwarded to the API.

Returns

  • Completion | Iterator[StreamChunk]: Completion if stream=False, Iterator[StreamChunk] if stream=True.

count_tokens​

count_tokens(*, model: str, prompt: str) -> dict[str, Any]

Count a prompt's tokens with a model's own tokenizer, and that model's context length.

Parameters

  • model (str): The model to tokenize against.
  • prompt (str): The text to count.

Returns

  • dict[str, Any]: count (the prompt's tokens) and max_model_len (the model's context window).

embedding​

embedding(*, model: str, input: str | list[str], **kwargs) -> EmbeddingResponse

Create embeddings.

Parameters

  • model (str): The model ID to use.
  • input (str | list[str]): Text or list of texts to embed.
  • **kwargs (Any): Additional parameters forwarded to the API.

generate​

generate(*, model: str, prompt: str, type: str | None = None, **kwargs) -> dict[str, Any]

Submit a generic generation job (umbrella endpoint).

Equivalent to the type-specific /ai-gateway/generations/{type} POSTs but the gateway dispatches by type.

Parameters

  • model (str): The generation model ID.
  • prompt (str): The text prompt.
  • type (str | None, optional): Generation type (image, video, or music).
  • **kwargs (Any): Additional generation parameters forwarded to the gateway.

generation_status​

generation_status(*, job_id: str) -> GenerationJob

Check the status of an async generation job.

image_generation​

image_generation(*, model: str, prompt: str, n: int = 1, size: str = "1024x1024", quality: str = "standard", response_format: str = "url", **kwargs) -> ImageGenerationResponse

Generate images from a text prompt.

list_speech_voices​

list_speech_voices() -> dict[str, Any]

List available TTS voices across providers.

moderation​

moderation(*, model: str, input: str | list[str], **kwargs) -> ModerationResponse

Check text for content policy violations.

music_generation​

music_generation(*, model: str, prompt: str, duration: int = 30, **kwargs) -> GenerationJob

Submit a music generation job (async).

rerank​

rerank(*, model: str, query: str, documents: list[str | dict[str, Any]], top_n: int | None = None, return_documents: bool = True, **kwargs) -> RerankResponse

Rerank documents by relevance to a query.

speech​

speech(*, model: str, input: str, voice: str = "alloy", response_format: str = "mp3", speed: float = 1.0, **kwargs) -> SpeechResponse

Generate speech audio from text (TTS).

transcription​

transcription(*, model: str, file: Any, filename: str = "audio.mp3", language: str | None = None, prompt: str | None = None, response_format: str = "json", temperature: float = 0.0, **kwargs) -> TranscriptionResponse

Transcribe audio to text (STT).

translation​

translation(*, model: str, file: Any, filename: str = "audio.mp3", **kwargs) -> TranscriptionResponse

Translate audio to English text (STT translation).

video_generation​

video_generation(*, model: str, prompt: str, duration: int = 5, resolution: str = "1080p", **kwargs) -> GenerationJob

Submit a video generation job (async).