Type

Results for “speech to text” — 17 073

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

Model Models & platforms Self-hosted

ekacare/parrotlet-e

huggingface.co

Модель feature-extraction · лицензия cc-by-sa-4.0

Model Models & platforms Self-hosted

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

Model Models & platforms Self-hosted

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

Model Models & platforms Self-hosted

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...

Model Models & platforms Self-hosted

OpenAI: GPT-4o

openrouter.ai

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...

Model Models & platforms Self-hosted

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

Model Models & platforms Self-hosted

namdp-ptit/ViDense

huggingface.co

Модель sentence-similarity · лицензия apache-2.0

Model Models & platforms Self-hosted

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

Model Models & platforms Self-hosted

Meta: Muse Spark 1.1

openrouter.ai

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...

Model Models & platforms Self-hosted

Модель token-classification · лицензия mit

Model Models & platforms Self-hosted

Submit a site to the catalog

Just send the link — we will work out the rest.

We will review what you send and add it to the catalog if it fits.

Not sure how to implement it? We can help

Tell us about your task — we will pick the tools and suggest where to start.

0 / 5000
Verification code

Fields marked with an asterisk are required. Your data is used only to reply.