Google "We Have No Moat, And Neither Does OpenAI"
AI competition statement
AI competition statement
A Stage Review of Instruction Tuning
LLM‑RL‑Visualized (EN) | LLM‑RL‑Visualized (中文) - 100+ LLM / RL Algorithm Maps📚.
ICML2022-Welcome to the "Big Model" Era: Techniques and Systems to Train and Serve Bigger Models
A Visual Guide to Mamba and State Space Models
CS324 - Large Language Models
ChatGPT Prompt Engineering
Recent Advances on Foundation Models.
high quality and educational materials you don't want to miss.
high quality and educational videos you don't want to miss.
My favorite!
TensorZero is an open-source framework for building production-grade LLM applications. It unifies an LLM gateway, observability, optimization, evaluations, and experimentation.
An all-in-one LLM Agent platform with your private data and knowledge, delivers your production-ready AI Agents on Day 1.
Deploy, manage, optimize any model at scale across any environment from cloud to edge. Let's you go from python notebook to inferencing in minutes.
A bilingual Chinese-English knowledge extraction model with knowledge graphs and natural language processing technologies.
A programming language for LLM interaction with support for typed prompting, control flow, constraints, and tools.
A paid product for detecting toxicity, hallucination, prompt injection, etc.
A paid product for testing and improving prompts.
A Python library for making chatbot interfaces.
Playground for devs to finetune & deploy LLMs
a toolkit for deploying and serving Large Language Models (LLMs).
TRL is a full stack library where we provide a set of tools to train transformer language models with Reinforcement Learning, from the Supervised Fine-tuning step (SFT), Reward Modeling step (RM) to the Proximal Policy Optimization (PPO) step.
</details>
Kimi-VL-A3B
InternVL-2|6|14|26
InternLM-Math-7B|20B
StarCoder2-3|7|15B
minicpm-2b-65d48bf958302b9fd25b698f)
OmniLLM-12B
CogVLM2-19B
</details>
Baichuan2-7|13B
</details>
Yi1.5-6|9|34B
</details>
</details>
</details>
OPT-1.3|6.7|13|30|66B
Qwen1.5-0.5B|1.8B|4B|7B|14B|32B|72B|110B|MoE-A2.7B
DeepSeek-MoE-16B
DeepSeek-Coder-1.3|6.7|7|33B
a benchmark dataset testing AI's ability to reason about visual commonsense through images that defy normal expectations.
a benchmark designed to assess the performance of multimodal web agents on realistic visually grounded tasks.
a large-scale question-answering benchmark focused on real-world financial data, integrating both tabular and textual information.
a biomedical question-answering benchmark designed for answering research-related questions using PubMed abstracts.
a multimodal question-answering benchmark designed to evaluate AI models' cognitive ability to understand human beliefs and goals.
a benchmark that evaluates large language models' ability to answer medical questions across multiple languages.
a comprehensive benchmarking platform designed to evaluate large models' mathematical abilities across 20 fields and nearly 30,000 math problems.