wren93/tuna

github.com
Visit site

Tuna: Taming Unified Visual Representations for Native Unified Multimodal Models [код обещают тут] Ранее мы много обсуждали мультимодальную генерацию с точки зрения: - Архитектуры : учить ли голову поверх LLM/VLM или делать unified backbone; - Представления данных : дискретное или непрерывное кодирование для картинок и текстов - Визуальных энкодеров : обычно для дискриминативных и генеративных задач используют разные (SigLip/VAE), но, например, Show-o2 ( статья, разбор )

Description

Specs

Type Tool
SectionImages
Site languageen

Found in sources

Similar in «Images»

Submit a site to the catalog

Just send the link — we will work out the rest.

We will review what you send and add it to the catalog if it fits.

Not sure how to implement it? We can help

Tell us about your task — we will pick the tools and suggest where to start.

0 / 5000
Verification code

Fields marked with an asterisk are required. Your data is used only to reply.