Skip to main content
POST
Audio To Text

Authorizations

Authorization
string
header
required

Bearer authentication header of the form Bearer <token>, where <token> is your auth token.

Body

multipart/form-data
audio
file
required

Uploaded audio file to be transcribed.

model_id
string
default:""

Hugging Face model ID used for transcription.

return_timestamps
string
default:true

Return timestamps for the transcribed text. Supported values: 'sentence', 'word', or a string boolean ('true' or 'false'). Default is 'true' ('sentence'). 'false' means no timestamps. 'word' means word-based timestamps.

metadata
string
default:{}

Additional job information to be passed to the pipeline.

Response

Successful Response

Response model for text generation.

text
string
required

The generated text.

chunks
Chunk · object[]
required

The generated text chunks.

Last modified on May 18, 2026