Type

Results for “speech to text” — 17 073

ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47B active per token. It is trained jointly on text and image data...

Model Models & platforms Self-hosted

KGESH/nsfw-bge-m3-v1

huggingface.co

Модель feature-extraction

Model Models & platforms Self-hosted

cstr/bge-m3-GGUF

huggingface.co

Модель feature-extraction · лицензия mit

Model Models & platforms Self-hosted

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

Model Models & platforms Self-hosted

Z.ai: GLM 5V Turbo

openrouter.ai

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

Model Models & platforms Self-hosted

Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math....

Model Models & platforms Self-hosted

Модель feature-extraction · лицензия apache-2.0

Model Models & platforms Self-hosted

cstr/jina-v5-nano-GGUF

huggingface.co

Модель feature-extraction · лицензия cc-by-nc-4.0

Model Models & platforms Self-hosted

Модель feature-extraction · лицензия apache-2.0

Model Models & platforms Self-hosted

Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video. The Instruct model targets general vision-language use (VQA, document parsing, chart/table...

Model Models & platforms Self-hosted

OpenAI: GPT-5 Image Mini

openrouter.ai

GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text...

Model Models & platforms Self-hosted

Модель sentence-similarity · лицензия apache-2.0

Model Models & platforms Self-hosted

Модель feature-extraction · лицензия apache-2.0

Model Models & platforms Self-hosted

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...

Model Models & platforms Self-hosted

cstr/f2llm-v2-80m-GGUF

huggingface.co

Модель feature-extraction · лицензия apache-2.0

Model Models & platforms Self-hosted

cstr/f2llm-v2-0.6b-GGUF

huggingface.co

Модель feature-extraction · лицензия apache-2.0

Model Models & platforms Self-hosted

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...

Model Models & platforms Self-hosted

Модель feature-extraction · лицензия apache-2.0

Model Models & platforms Self-hosted

Submit a site to the catalog

Just send the link — we will work out the rest.

We will review what you send and add it to the catalog if it fits.

Not sure how to implement it? We can help

Tell us about your task — we will pick the tools and suggest where to start.

0 / 5000
Verification code

Fields marked with an asterisk are required. Your data is used only to reply.