AI Inference
The AI Inference resource is the OpenAI-compatible gateway for chat completions, text completions, embeddings, and audio. Reference a model by its vendor id; the gateway resolves the provider key and routes the call.
Access it as client.ai.inference on a Strongly client, or the same path on AsyncStrongly with await. All methods exist on both with identical signatures.
Methods
All methods
cancel_generation
cancel_generation(*, job_id: str) -> GenerationJob
Cancel an async generation job.
chat_completion
chat_completion(*, model: str, messages: Sequence[dict[str, Any] | ChatMessage], stream: bool = False, max_tokens: int | None = None, temperature: float = 0.7, top_p: float = 1.0, stop: str | list[str] | None = None, **kwargs) -> ChatCompletion | Iterator[StreamChunk]
Send a chat completion request.
Parameters
model(str): The model ID to use.messages(Sequence[dict[str, Any] | ChatMessage]): List of chat messages.stream(bool, optional): If True, return an iterator of StreamChunk objects.max_tokens(int | None, optional): Maximum tokens to generate.temperature(float, optional): Sampling temperature.top_p(float, optional): Nucleus sampling parameter.stop(str | list[str] | None, optional): Stop sequence(s).**kwargs(Any): Additional parameters forwarded to the API.
Returns
ChatCompletion | Iterator[StreamChunk]: ChatCompletion if stream=False, Iterator[StreamChunk] if stream=True.
completion
completion(*, model: str, prompt: str, stream: bool = False, max_tokens: int | None = None, temperature: float = 0.7, **kwargs) -> Completion | Iterator[StreamChunk]
Send a text completion request.
Parameters
model(str): The model ID to use.prompt(str): The text prompt.stream(bool, optional): If True, return an iterator of StreamChunk objects.max_tokens(int | None, optional): Maximum tokens to generate.temperature(float, optional): Sampling temperature.**kwargs(Any): Additional parameters forwarded to the API.
Returns
Completion | Iterator[StreamChunk]: Completion if stream=False, Iterator[StreamChunk] if stream=True.
count_tokens
count_tokens(*, model: str, prompt: str) -> dict[str, Any]
Count a prompt's tokens with a model's own tokenizer, and that model's context length.
Parameters
model(str): The model to tokenize against.prompt(str): The text to count.
Returns
dict[str, Any]:count(the prompt's tokens) andmax_model_len(the model's context window).
embedding
embedding(*, model: str, input: str | list[str], **kwargs) -> EmbeddingResponse
Create embeddings.
Parameters
model(str): The model ID to use.input(str | list[str]): Text or list of texts to embed.**kwargs(Any): Additional parameters forwarded to the API.
generate
generate(*, model: str, prompt: str, type: str | None = None, **kwargs) -> dict[str, Any]
Submit a generic generation job (umbrella endpoint).
Equivalent to the type-specific /ai-gateway/generations/{type} POSTs but the
gateway dispatches by type.
Parameters
model(str): The generation model ID.prompt(str): The text prompt.type(str | None, optional): Generation type (image,video, ormusic).**kwargs(Any): Additional generation parameters forwarded to the gateway.
generation_status
generation_status(*, job_id: str) -> GenerationJob
Check the status of an async generation job.
image_generation
image_generation(*, model: str, prompt: str, n: int = 1, size: str = "1024x1024", quality: str = "standard", response_format: str = "url", **kwargs) -> ImageGenerationResponse
Generate images from a text prompt.
list_speech_voices
list_speech_voices() -> dict[str, Any]
List available TTS voices across providers.
moderation
moderation(*, model: str, input: str | list[str], **kwargs) -> ModerationResponse
Check text for content policy violations.
music_generation
music_generation(*, model: str, prompt: str, duration: int = 30, **kwargs) -> GenerationJob
Submit a music generation job (async).
rerank
rerank(*, model: str, query: str, documents: list[str | dict[str, Any]], top_n: int | None = None, return_documents: bool = True, **kwargs) -> RerankResponse
Rerank documents by relevance to a query.
speech
speech(*, model: str, input: str, voice: str = "alloy", response_format: str = "mp3", speed: float = 1.0, **kwargs) -> SpeechResponse
Generate speech audio from text (TTS).
transcription
transcription(*, model: str, file: Any, filename: str = "audio.mp3", language: str | None = None, prompt: str | None = None, response_format: str = "json", temperature: float = 0.0, **kwargs) -> TranscriptionResponse
Transcribe audio to text (STT).
translation
translation(*, model: str, file: Any, filename: str = "audio.mp3", **kwargs) -> TranscriptionResponse
Translate audio to English text (STT translation).
video_generation
video_generation(*, model: str, prompt: str, duration: int = 5, resolution: str = "1080p", **kwargs) -> GenerationJob
Submit a video generation job (async).