tscosj

Submissions and activity by tscosj on LLM Atlas.

Mistral 7B

Compact 7B model outperforming Llama 2 13B with grouped-query and sliding window attention

model

llama.cpp

Efficient C/C++ LLM inference — run models locally on CPU and consumer GPUs with quantization

repo

LangChain

Framework for building LLM applications — chains, agents, retrieval, and memory

repo

LlamaIndex

Data framework connecting custom data sources to LLMs with indexing, retrieval, and query engines

repo

vLLM

High-throughput LLM serving engine with PagedAttention for efficient memory management

repo

Ollama

Run LLMs locally with a single command — supports Llama, Mistral, Gemma, and dozens more

tool

LiteLLM

Universal API proxy — call 100+ LLM providers with one interface, with fallbacks and load balancing

tool

OpenAI Python SDK

Official Python client for the OpenAI API — GPT, DALL-E, Whisper, and embeddings

repo

MLflow

Open source platform for ML lifecycle — experiment tracking, model registry, and deployment

tool

Hugging Face

The leading platform for ML models, datasets, and demos — hub for the open-source AI community

tool

Weights & Biases

ML experiment tracking, visualization, and collaboration platform used by top research labs

tool

MMLU Benchmark

Massive Multitask Language Understanding — 57-subject benchmark covering STEM, humanities, and more

benchmark

HumanEval

164 handwritten Python problems for evaluating LLM code generation ability

benchmark

Common Crawl

Petabytes of raw web data used to train most large language models — the internet as a dataset

dataset

Hugging Face Spaces

Host and share ML demos with Gradio and Streamlit — free GPU access for interactive applications

demo

r/LocalLLaMA

Reddit community for running LLMs locally — hardware guides, model comparisons, quantization tips

discussion

r/MachineLearning

Reddit's largest ML community — paper discussions, industry news, career advice, and research

discussion

ChatGPT

OpenAI's flagship chat — GPT-4o with voice, image, file, and web search capabilities

chat

Claude

Anthropic's assistant — long context, code, analysis, and artifacts with Projects support

chat

Gemini

Google DeepMind's multimodal assistant with 1M context, Deep Research, and Workspace integration

chat

Meta AI

Meta's AI assistant powered by Llama — available on WhatsApp, Instagram, Facebook, and web

chat

Microsoft Copilot

Microsoft's AI assistant built on GPT-4 with Bing search, image generation, and Office integration

chat

Perplexity

AI-powered answer engine with real-time web search and cited sources — great for research

chat

Le Chat

Mistral AI's chat interface — fast inference on Mistral Large, Pixtral, and open models

chat

Grok

xAI's assistant with real-time X/Twitter data access and unfiltered responses

chat

DeepSeek Chat

DeepSeek's chat with extended thinking mode — strong on math, code, and reasoning tasks

chat

HuggingChat

Hugging Face's open-source chat — switch between Llama, Mistral, Qwen, and community models

chat

Poe

Quora's multi-model hub — access ChatGPT, Claude, Gemini, and custom bots in one place

chat

You.com

AI search and chat with web access, code interpreter, and document analysis

chat

Pi

Inflection AI's personal assistant focused on emotional intelligence and ongoing conversation

chat

Cohere Coral

Cohere's enterprise chat grounded in your own documents — RAG-native assistant

chat

Character.AI

Create and chat with AI personas — millions of community-built characters for roleplay and learning

chat

Kagi Assistant

Privacy-first AI assistant from Kagi with multiple model choices and no ads or tracking

chat

Phind

Code-focused AI search and chat — understands your codebase, stack traces, and developer questions

chat

GitHub Copilot Chat

GitHub's AI assistant embedded in github.com — explain, fix, and generate code with repo context

chat

GroqCloud Playground

Chat at extreme speed on Groq LPUs — try Llama, Mixtral, Gemma at 500+ tokens/sec

chat

OpenRouter Chat

One interface for 200+ models — compare Claude, GPT, Gemini, open-source, and experimental models

chat

NVIDIA NIM

Try and deploy NVIDIA-optimized LLMs — Llama, Mistral, and community models with one-click API access

chat

Amazon Q

Amazon's AI assistant for AWS and business — grounded in your data, code, and cloud infrastructure

chat

Kimi

Moonshot AI's assistant with 1M token context — strong at long documents and Chinese/English tasks

chat

Doubao

ByteDance's AI chat — largest user base in China, strong multimodal and creative writing capabilities

chat

ChatGLM

Zhipu AI's bilingual LLM — open-source GLM-4 with strong Chinese and code performance

chat

Tongyi Qianwen

Alibaba's Qwen-powered assistant — strong multilingually, long context, free to use

chat

LM Studio

Desktop app to run any GGUF model locally — one-click downloads, chat UI, and local server mode

chat

Jan

Open-source offline-first desktop chat — runs Llama, Mistral, and custom models completely on-device

chat

Open WebUI

Self-hosted Ollama and OpenAI-compatible chat UI — full featured, RAG, multimodal, and pipelines

chat

Msty

Local and cloud model chat with side-by-side comparison, branching conversations, and knowledge stacks

chat

Venice

Privacy-preserving AI chat — conversations never leave your device, no logs, uncensored open models

chat

Cerebras Systems

Wafer-scale AI processors. WSE-3: 4T transistors, 900K cores, 125 PFLOPS. Raised $1B Series H led by Tiger Global. IPO expected 2026.

company

MatX

LLM chip with splittable systolic array. SRAM-first + HBM for long-context. Targets training, RL, and inference. Raised $500M Series B.

company

Ayar Labs

In-package optical interconnects for AI. Co-packaged optics with multi-wavelength light sources, terabit bandwidth mm-to-km. Raised $500M Series E.

company

Rebellions

AI processors for large-scale inference — LLMs, MoE, multimodal. Four-chiplet SoC with UCIe-Advanced mesh, HBM3E. Raised $400M pre-IPO.

company

SambaNova Systems

Reconfigurable dataflow unit (RDU) architecture for inference. Streaming pipeline replaces kernel-by-kernel execution. 5th-gen chip links 256 accelerators. Raised $350M Series E.

company

Ricursive Intelligence

AI models to automate all stages of chip design and verification. Recursive feedback loop: AI designs AI chips that enable better AI. Raised $300M Series A.

company

Axelera AI

Inference acceleration for GenAI and computer vision. SRAM-based digital in-memory computing with RISC-V dataflow architecture. Raised $250M+.

company

Positron AI

Energy-efficient hardware for long-context transformer inference. FPGA server with 93%+ memory bandwidth utilization, supports trillion-parameter models. Raised $230M Series B.

company

Kandou AI

High-speed signaling and SerDes for AI chip-to-chip interconnects. Chord multi-wire signaling enables better signal integrity at lower power than PAM-4. Raised $225M.

company

Olix

Optical tensor processing unit for inference. SRAM-based architecture with optical digital processor for bit-perfect logic. Raised $220M Series A.

company

Eridu

Network switch for AI data centers. Unifies silicon, optics, packaging, and software. Supports scale-out of 100K+ GPUs. Raised $200M Series A from stealth.

company

Upscale AI

Ultra-low latency AI networking silicon, systems, and software. Open standards (ESUN, UALink, Ultra Ethernet). Unifies GPUs, accelerators, memory into single AI engine. Raised $200M Series A.

company

Taalas

Automated flow for implementing any AI model in silicon. Unifies storage and compute at DRAM-level density. First product: hard-wired Llama 3.1 8B in silicon. Raised $169M.

company

Rapidus

Pure-play foundry building 2nm fab in Japan. Mass production expected 2027. Raised ~$1.7B from Japan govt + Canon, Sony, SoftBank, NTT, Fujitsu.

company

Neurophos

Optical processing unit with 1M+ metamaterial elements on single chip. Analog systolic array. 0.47 ExaOPS FP4 at 235 TOPS/W. GPU drop-in replacement. Raised $110M Series A.

company

Efficient Computer

Reconfigurable dataflow processor. Programs expressed as instruction circuits laid out spatially across simple processor array. Eliminates unnecessary data movement. Raised $60M Series A.

company

ChipAgents

Agentic AI for chip design. AI agents automate RTL design, debugging, and verification. Natural language to design specs, auto-generates Verilog, autonomous verification. Raised $50M Series A1.

company

Normal Computing

AI-powered EDA platform using auto-formalization (LLMs + formal logic). Designs thermodynamic computing silicon. Taped out first chip for multi-modal diffusion GenAI inference. Raised $50M.

company

Eliyan Corporation

Chiplet interconnect tech for multi-die AI architectures. NuLink PHY with simultaneous bidirectional signaling. Memory and I/O chiplets for 1.6-12.8 Tbps links. Raised $50M.

company

Mesh Optical Technologies

High-volume optical manufacturing for AI workloads. 1.6 Tb/s optical transceiver with flip-chip die bonding. Focus on scalable optical packaging. Raised $50M+ Series A.

company

Vertical Compute

Vertically integrated memory and compute chiplets for AI accelerators. High-aspect-ratio 3D structure with vertical data lanes on top of compute units. imec spin-off. Raised ~$43M seed.

company

Xscape Photonics

Programmable multi-wavelength photonics for AI data center fabrics. WDM fabric roadmap to 128 colors for ultra-fast optical data transmission. Raised $37M Series A extension.

company

Quadric

General-purpose NPU processor IP. C++ programmable hybrid Von Neumann + 2D SIMD architecture. Scalable 1-864 TOPS. Targets edge LLM, automotive, vision. Raised $30M Series C.

company

Claros

Chip-to-grid power management for AI data centers. Integrated voltage regulator delivers power directly to processing units. DC-native distribution eliminates AC/DC conversion losses. Raised $30M seed.

company

Great Sky

AI chips combining superconducting computation, optical communication, and brain-inspired architecture. Co-located memory and processing enables on-device adaptation without retraining. Raised $14M seed.

company

Cursor raises $900M Series C at $9B valuation

The AI-native code editor built on VS Code reports 1M+ daily active developers. Backed by a]16z, Thrive, and Spark. Cursor Tab and multi-file editing drive adoption among professional developers.

news

Why are AI agents lying, cheating and coordinating?

A lot has been written about the incidents of the last few months in which AI agents misbehaved in serious ways. They took actions that would be considered as crimes if a human took them, escaped their containment to cheat on assigned tasks while attempting to evade detection,…

news

OpenAI agents attacked RubyGems back in May

OpenAI agents carried out an undisclosed attack on RubyGems is a new bombshell report from Spencer Kitts, Thomas Larsen, and Sydney Von Arx - three of the four authors of the report on the agent attack on disused wikis (

news

vllm v0.29.1rc0

[watermarking] Dual-key gumbel-max watermarking for speculative decod…

release

Claude is only available to people over 18 years

Claude, our consumer product, is only available to people over 18 years. You’ll need to confirm you’re 18 or over while setting up an account. When we detect signals that you may be under 18, we'll ask you to verify your age before you can continue using Claude.

news

The Gradient of Generative AI Release: Methods and Considerations

As increasingly powerful generative AI systems are developed, the release method greatly varies. We propose a framework to assess six levels of access to generative AI systems: fully closed; gradual or staged access; hosted access; cloud-based or API access; downloadable access;…

paper

ollama v0.34.0

Use Ollama models in ChatGPT Desktop Ollama models can now be used directly in ChatGPT Desktop, so you can keep your existing workflow while running open models. Setup is available from the Ollama app on MacOS. This release also improves structured output performance on Apple…

release

Some thoughts on the Navier–Stokes Millennium Prize Problem

On the Navier–Stokes Millennium Prize Problem introduces an impressive result from OpenAI, who used an unreleased model to produce a resolution to the Navier–Stokes existence and smoothness problem , one of the seven Millennium Prize Problems that have been subject to a…

news

llama.cpp b10944

sycl : Fix get mem error ( #28227 ) fix for unsupport zes API optimize the code adjust the log level rm unused head files Update docs/backend/SYCL.md Co-authored-by: Titaniumtown titaniumtown@proton.me fix the error to detect level zero SDK/dev package, stop build after detect…

release

llama.cpp b10946

ggml-cpu(s390x): guard VXE-only repack helpers ( #28775 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47190488 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…

release

llama.cpp b10947

models : guard the expert FFN size fallback in nemotron-h against a zero divisor ( #28779 ) The NextN/MTP tail loop derives the expert FFN size as n_ff/n_expert_used when expert_feed_forward_length gives nothing for the layer. Both values come from per-layer arrays that…

release

llama.cpp b10948

tests : exclude HY_V4 from WebGPU test-llama-archs tests ( #28855 ) Co-authored-by: Stanisław Szymczyk sszymczy@gmail.com Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47194535 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple…

release

llama.cpp b10950

ggml-cuda: fallback to F32 on device without BF16 hardware acceleration ( #28846 ) ggml-cuda: fallback to F32 on device without BF16 hardware acceleration: (Nvidia >= AMPERE, AMD >= RDNA3 or = CDNA) apply logic to NVIDIA as well Co-authored-by: Johannes Gäßler johannesg@5d6.de…

release

AI is not a normal technology

This is an addendum to my original Meditations on AI post, and a response to a few other articles from last year.

news

llama.cpp b10951

common : move llama_n_rs_seq to before llama_decode ( #28749 ) This commit moves the llama_n_rs_seq function call to before the llama_decode call and returns directly if the check is true, removing the setting of res and the goto statement. The motivation for this change is to…

release

Who gets to define the rules for AI?

Artificial intelligence is remaking the world we live in. Within a generation, the way we discover medicine, manage power grids, and secure our national infrastructure will be completely transformed. Many people already realize this and are working to build that future…

news

llama.cpp b10952

sycl : fix oneDNN scratchpad breaking the pool free order ( #28704 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47287268 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel…

release

The Three AI Pills

Sincere disagreements about AI are usually disagreements about future AI capabilities. There are roughly four positions people take. Two are reasonable. Two are not. I distinguish these via the Thr…

news

llama.cpp b10955

ggml-cpu : disable PCH and fix CACHE_LINE_SIZE ambiguity to fix heap corruption ( #28882 ) Disable the ggml-cpu precompiled header and remove the std::hardware_destructive_interference_size branch from CACHE_LINE_SIZE. The PCH force-includes ggml-impl.h before ops.h, which pulls…

release

Apple's Siri AI Can Be Swapped Out for Claude, ChatGPT, Code Shows

Code sleuth "pdfu" has uncovered iOS 27 and macOS Golden Gate private frameworks that show Apple has designed its new Siri architecture to work with third-party AI models at what appears to be a surprisingly deep level. One mechanism called Model Delegation allows Claude to…

news

What a time to be alive – rouge AI agents attack RubyGems.org

Today Reuters and the Wall Street Journal both reported about rogue AI agents at OpenAI attacking RubyGems.org. https://www.rubyhack.ai/ has an amazing writeup, and you should read it. I just wanted to make a quick post about it because it’s wild. TL;DR: It seems like OpenAI…

news

llama.cpp b10956

sycl: rfc: Use radix select for top_k ( #28670 ) sycl: GPU-resident TOP_K for large k, parallelised over the device The SYCL backend refused GGML_OP_TOP_K above k = 32 and let it fall back to the CPU, a backend round-trip per call. The limit was not conservatism: the scan-merge…

release

The AI job market in 2026

The AI labor market in September 2026 from LinkedIn, WEF, Stanford, PwC and Bain data: the most in-demand AI roles, salaries in the US and Europe, what grows and what fades.

news

llama.cpp b10964

llama.cpp : bump version to 0.4.1 ( #28900 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47376843 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework…

release

self improving agent

Today's AI systems have human-designed, fixed architectures and cannot autonomously and continuously improve themselves. The advance of AI could itself be automated. If done safely, that would accelerate AI development and allow us to reap its benefits much sooner. Meta-learning…

paper

The Eureka Machine

An AI expert reveals the transformative potential of artificial intelligence to usher in a new era of scientific discovery.“There is a lot being written ab...

tool

For AI leaders Doom is a form of hype

Why AI doom rhetoric from Anthropic, OpenAI and other tech leaders functions as hype, regulatory strategy, and a distraction from present harms.

news

Hacking AI customer service agents

As AI agents are deployed to automate more tasks, they become more capable. And as the famous quote goes: "With great power comes great responsibility." Assuming that humans in the loop can mitigate t...

news

llama.cpp b10969: ci : add ubuntu-cuda builds to release (#28186)

release : add ubuntu-cuda build job (12.8/13.3, x64+arm64) Add GCC 14 for CUDA arm64 builds in CI Eplicit bash Install git for CCCL fetch Install git before we clone/checkout Match CI names for WIndows Whitelist llama.cpp repo to git Use $GITHUB_WORKSPACE Also ship dependent…

release

llama.cpp v0.4.1

Overview llama.cpp 0.4.1 adds Maple 20B-A1B, Tencent Hy 4, and Spark2.5 support, improves JSON schema handling, chat parsing, logging, and server child-process management, and updates ggml to v0.24.0. API changes Changed llama_sampler_chain_n() to return int32_t instead of int (…

release

llama.cpp b10970

HIP: fattn-mma: use fp32 accumulation on MFMA devices ( #28576 ) use fp32 accumulators in fattn-mma on CDNA Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47450174 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64,…

release

Anthropic is in regulatory-capture financial loop

The AI safety nonprofits around Anthropic promote an Anthropic-aligned Doomer narrative, whether Tarbell Fellows in popular media, or METR in evaluations. All rely on Moskovitz's Anthropic stock worth $7 billion, which the parent funding org sits on. This stock increases in…

news

Show HN: Sunk Cost – How long until a local LLM rig pays for itself?

I kept hearing "just buy a Mac and run models locally, it pays for itself" and wanted to check. Sunk Cost takes a machine, a model and how many tokens you use a day, and works out how long the hardware takes to pay back against renting the same model by the token. Obviously…

news

rare personal blogpost

I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life. I’ve written often about these incredible benefits: I believe that AI could cure most major diseases in the next 5–10 years, greatly accelerate economic growth…

tool

opt in/out

Via [AINews] AEF-1 standard emerges for Third Party Evaluators, as Xai, OpenAI, and Anthropic all cosign

news

MMLU

Multiple-choice questions across 57 subjects, from elementary math to law and medicine. Measures broad world knowledge and problem solving.

benchmark

MMLU-Pro

A harder MMLU variant with 10 answer options and more reasoning-heavy questions, designed to reduce prompt sensitivity.

benchmark

MMLU-Redux

A cleaned subset of MMLU with flawed or ambiguous questions removed, plus a harder MMLU-Pro style split.

benchmark

GPQA Diamond

Graduate-level physics, chemistry and biology questions written to be Google-proof. Diamond is the hardest verified subset.

benchmark

SuperGPQA

Graduate-level questions across 285 disciplines, billed as the largest knowledge benchmark of its kind.

benchmark

Humanity's Last Exam

Thousands of expert-written questions across dozens of fields, designed to be at the frontier of what models can answer.

benchmark

ARC-AGI

Novel visual grid puzzles that require inferring a rule from a few examples, testing fluid intelligence rather than recall.

benchmark

ARC-AGI-2

A harder successor to ARC-AGI with puzzles tuned so that brute-force or memorisation does not help.

benchmark

MuSR

Multi-step soft-reasoning problems written so they cannot be solved by pattern matching a single sentence.

benchmark

AGIEval

Standardised exams (SAT, GRE, LSAT, Chinese gaokao and more) used to test human-level problem solving.

benchmark

WinoGrande

Pronoun-resolution puzzles that require commonsense and careful reading rather than world knowledge.

benchmark

ZebraLogic

Logic-grid (Zebra) puzzles of controlled size, measuring structured constraint reasoning and search.

benchmark

C-Eval

A Chinese-language multi-subject knowledge benchmark spanning STEM, humanities and social science.

benchmark

CMMLU

A Chinese multi-task benchmark covering 67 topics, with an emphasis on China-specific knowledge.

benchmark

GSM8K

Grade-school math word problems requiring multi-step arithmetic reasoning.

benchmark

MATH

Competition mathematics problems with full worked solutions, spanning algebra through calculus.

benchmark

AIME

Problems from the American Invitational Mathematics Examination, used as a hard, contamination-resistant math set.

benchmark

FrontierMath

Original, unpublished research mathematics problems vetted by professional mathematicians.

benchmark

MathArena

Continuously refreshed evaluation on the latest live math competitions, released as they happen.

benchmark

OlympiadBench

Olympiad-level math and physics problems with images and text, in English and Chinese.

benchmark

HumanEval

164 hand-written Python programming problems scored by running unit tests. The classic code-generation benchmark.

benchmark

MBPP

Mostly Basic Python Problems: around 1,000 crowd-sourced entry-level programming tasks with test cases.

benchmark

EvalPlus (HumanEval+ / MBPP+)

HumanEval and MBPP hardened with far more test cases, exposing solutions that only pass the originals by luck.

benchmark

BigCodeBench

Realistic, practical programming tasks that require composing multiple libraries and function calls.

benchmark

LiveCodeBench

Contamination-resistant competitive programming problems collected after a model's training cut-off.

benchmark

SWE-bench Verified

Real GitHub issues that a model must fix in a repository; Verified is the human-validated 500-issue subset.

benchmark

Aider Polyglot

Exercises in multiple languages where the model must edit existing code, scored including its formatting and diff errors.

benchmark

RepoBench

Repository-level code completion across files and languages, testing cross-file context use.

benchmark

DS-1000

Data-science coding problems covering NumPy, pandas, PyTorch and other libraries, with executable checks.

benchmark

Spider

Text-to-SQL across many databases and schemas, scored on whether the generated query returns the right rows.

benchmark

BIRD

A larger, messier text-to-SQL benchmark with dirty data and external knowledge, closer to real databases.

benchmark

CRUXEval

Short Python functions where the model must predict the output or the input, testing code execution reasoning.

benchmark

DROP

Discrete reasoning over paragraphs: answer questions that require arithmetic, counting, and sorting over text.

benchmark

BIG-Bench Hard

23 BIG-Bench tasks where models initially performed below the average human rater, focused on multi-step reasoning.

benchmark

ARC-Challenge

Grade-school science questions that require reasoning rather than recall, from the AI2 Reasoning Challenge.

benchmark

HellaSwag

Commonsense sentence completion: choose the most plausible continuation of a described situation.

benchmark

TruthfulQA

Questions designed to elicit common misconceptions, testing whether a model reproduces falsehoods.

benchmark

SimpleQA

Short factual questions with a single verifiable answer, scored for both accuracy and hallucination rate.

benchmark

FACTS Grounding

Long-form answers judged for whether every claim is supported by the provided source document.

benchmark

HaluEval

A suite for detecting hallucinated content across QA, dialogue and summarisation tasks.

benchmark

IFEval

Verifiable instruction-following tasks (format, length, casing, keywords) checked programmatically.

benchmark

MT-Bench

Multi-turn conversation quality judged by a strong model acting as grader across eight categories.

benchmark

Chatbot Arena Elo

Crowd-sourced pairwise preference ranking from blind battles against other models.

benchmark

LiveBench

A contamination-resistant benchmark with a fresh set of questions released regularly and objective ground truths.

benchmark

WildBench

Real-world user prompts drawn from chats, graded against reference answers by an LLM judge.

benchmark

AlpacaEval

Head-to-head instruction-following judged by an LLM, reported as a win rate against a reference model.

benchmark

Arena-Hard

Hard prompts from Chatbot Arena, scored by an LLM judge against a fixed baseline for reproducibility.

benchmark

MixEval

Real user queries matched to a ground-truth benchmark, aiming to predict human preference cheaply.

benchmark

MMMU

College-level questions with images and diagrams across six disciplines, testing multimodal reasoning.

benchmark

MMMU-Pro

A tougher MMMU that embeds questions in screenshots and removes answer shortcuts.

benchmark

MathVista

Mathematical reasoning in visual contexts — charts, geometry, and figures — for multimodal models.

benchmark

MMBench

A broad multimodal benchmark covering perception, reasoning, and fine-grained visual ability.

benchmark

ChartQA

Questions about charts that need visual and logical reasoning, not just reading a number off an axis.

benchmark

DocVQA

Questions answered from scanned document images, testing OCR and layout understanding.

benchmark

AI2D

Multiple-choice questions about science diagrams, testing diagram understanding and reasoning.

benchmark

RealWorldQA

Everyday photographic questions about spatial and real-world understanding.

benchmark

Video-MME

Video question answering across short clips to long videos, with and without subtitles.

benchmark

Vibe-Eval

Hard visual prompts judged by humans and a model, aimed at day-to-day multimodal usefulness.

benchmark

τ-bench

Tool-and-agent tasks (retail, airline) where a model must follow policy across multi-step conversations and tool calls.

benchmark

τ²-bench

A successor to τ-bench with more domains and dual-control tasks where the user and agent both act.

benchmark

Terminal-Bench

End-to-end tasks completed in a real terminal sandbox, testing long-horizon agentic execution.

benchmark

GAIA

Real-world assistant questions that require browsing, tools, and multi-step reasoning to answer.

benchmark

OSWorld

Real computer tasks (files, apps, browser) executed on a live operating system and verified by state.

benchmark

WebArena

Web tasks in self-hosted replicas of real sites (shopping, forums, maps), scored by final state.

benchmark

ToolBench

Tool-use tasks over a large set of real APIs, testing planning, selection and call accuracy.

benchmark

AgentBench

An evaluation of models as agents across eight environments, from OS to games to web.

benchmark

MLE-bench

Kaggle machine-learning competitions an agent must complete end to end, scored by medal threshold.

benchmark

LongBench

A suite of long-context tasks (QA, summarization, code) across multiple lengths and languages.

benchmark

RULER

Synthetic long-context tasks at controllable lengths that reveal where effective context starts to break down.

benchmark

BABILong

Reasoning over very long texts built from bAbI tasks, stressing recall over millions of tokens.

benchmark

InfiniteBench

Tasks averaged over extremely long inputs (100k+ tokens), including retrieval and summarisation.

benchmark

Needle in a Haystack

A simple test of whether a model can find one fact placed at a chosen depth in a long context.

benchmark

HELMET

A broad long-context evaluation spanning retrieval, generation, and reasoning at multiple lengths.

benchmark

HarmBench

A standardised evaluation of how often harmful requests or jailbreaks succeed against a model.

benchmark

JailbreakBench

A reproducible benchmark of jailbreak attacks and defences with a public leaderboard.

benchmark

WMDP

A proxy measurement of hazardous knowledge in biosecurity, cybersecurity, and chemical domains.

benchmark

MedQA (USMLE)

Medical licensing exam questions, a common proxy for clinical knowledge.

benchmark

LegalBench

A collaboratively built set of legal reasoning tasks spanning many practice areas.

benchmark

FinanceBench

Open-book questions over real financial filings, testing retrieval and numerical grounding.

benchmark

llama.cpp b10976

ci : fix android release ( #28936 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47548995 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu…

release

llama.cpp b10977

ci: Bump CUDA Windows x64 builds to 13.4.1 ( #28930 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47583571 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…

release