Grok Speech to Text

Model Information

Grok Speech to Text (STT) is xAI's standalone speech recognition API for batch and real-time transcription. It supports multilingual transcription, word-level timestamps, speaker diarization, multichannel audio, keyterm prompting, Inverse Text Normalization, and real-time interim results over WebSocket.

Model ID

xai.grok-stt

Use this ID when making API calls to reference this model

Provider

SpaceXAI

Model Type

ASR

Accuracy Tier

premium

Release Date

April 17, 2026

Supported Languages

arcsdanlenfilfrdehiiditjakomkmsfaplptroruessvthtrvi
Automatic Language Detection: Yes

Streaming Transcription Languages

arcsdanlenfilfrdehiiditjakomkmsfaplptroruessvthtrvi
Performance & Cost

Cost

$0.10000/hour

$0.00003/second

Maximum File Size

500.00 MB

Features

Supported capabilities and functionalities

Core Features

Punctuation
Diarization
Streaming
Speaker Labels
Word Timestamps
Confidence Scores
Custom Vocabulary
Profanity Filtering
Noise Reduction
Voice Activity Detection

Subtitle Formats

SRT Support
VTT Support
Technical Specifications

Input/output formats and technical details

Subtitle Format Support

No subtitle formats supported

Supported Audio Encodings

WAVMP3OGGOpusFLACAACMP4M4AMKVPCMmulawalaw

Supported Sample Rates

8000 Hz16000 Hz22050 Hz24000 Hz44100 Hz48000 Hz