Attention Is All You Need
The foundational paper introducing the Transformer architecture that powers all modern LLMs
paperSubmissions and activity by tscosj on LLM Atlas.
The foundational paper introducing the Transformer architecture that powers all modern LLMs
paperBidirectional encoder representations that revolutionized NLP transfer learning
paperThe GPT-3 paper showing that scaling language models enables few-shot learning across tasks
paperParameter-efficient fine-tuning that reduces trainable parameters by orders of magnitude
paperPrompting LLMs to think step-by-step dramatically improves reasoning performance
paperTraining helpful, harmless, and honest AI systems using AI-generated feedback
paperMeta's open-source LLM family from 7B to 70B parameters, available for research and commercial use
modelCompact 7B model outperforming Llama 2 13B with grouped-query and sliding window attention
modelState-of-the-art ML library with thousands of pretrained models for text, vision, and audio
repoEfficient C/C++ LLM inference — run models locally on CPU and consumer GPUs with quantization
repoData framework connecting custom data sources to LLMs with indexing, retrieval, and query engines
repoUniversal API proxy — call 100+ LLM providers with one interface, with fallbacks and load balancing
toolOfficial Python client for the OpenAI API — GPT, DALL-E, Whisper, and embeddings
repoOpen source platform for ML lifecycle — experiment tracking, model registry, and deployment
toolThe leading platform for ML models, datasets, and demos — hub for the open-source AI community
toolML experiment tracking, visualization, and collaboration platform used by top research labs
toolMassive Multitask Language Understanding — 57-subject benchmark covering STEM, humanities, and more
benchmarkPetabytes of raw web data used to train most large language models — the internet as a dataset
datasetHost and share ML demos with Gradio and Streamlit — free GPU access for interactive applications
demoReddit community for running LLMs locally — hardware guides, model comparisons, quantization tips
discussionReddit's largest ML community — paper discussions, industry news, career advice, and research
discussionGoogle DeepMind's multimodal assistant with 1M context, Deep Research, and Workspace integration
chatMeta's AI assistant powered by Llama — available on WhatsApp, Instagram, Facebook, and web
chatMicrosoft's AI assistant built on GPT-4 with Bing search, image generation, and Office integration
chatAI-powered answer engine with real-time web search and cited sources — great for research
chatDeepSeek's chat with extended thinking mode — strong on math, code, and reasoning tasks
chatHugging Face's open-source chat — switch between Llama, Mistral, Qwen, and community models
chatCreate and chat with AI personas — millions of community-built characters for roleplay and learning
chatPrivacy-first AI assistant from Kagi with multiple model choices and no ads or tracking
chatCode-focused AI search and chat — understands your codebase, stack traces, and developer questions
chatGitHub's AI assistant embedded in github.com — explain, fix, and generate code with repo context
chatChat at extreme speed on Groq LPUs — try Llama, Mixtral, Gemma at 500+ tokens/sec
chatOne interface for 200+ models — compare Claude, GPT, Gemini, open-source, and experimental models
chatTry and deploy NVIDIA-optimized LLMs — Llama, Mistral, and community models with one-click API access
chatAmazon's AI assistant for AWS and business — grounded in your data, code, and cloud infrastructure
chatMoonshot AI's assistant with 1M token context — strong at long documents and Chinese/English tasks
chatByteDance's AI chat — largest user base in China, strong multimodal and creative writing capabilities
chatAlibaba's Qwen-powered assistant — strong multilingually, long context, free to use
chatDesktop app to run any GGUF model locally — one-click downloads, chat UI, and local server mode
chatOpen-source offline-first desktop chat — runs Llama, Mistral, and custom models completely on-device
chatSelf-hosted Ollama and OpenAI-compatible chat UI — full featured, RAG, multimodal, and pipelines
chatLocal and cloud model chat with side-by-side comparison, branching conversations, and knowledge stacks
chatPrivacy-preserving AI chat — conversations never leave your device, no logs, uncensored open models
chatWafer-scale AI processors. WSE-3: 4T transistors, 900K cores, 125 PFLOPS. Raised $1B Series H led by Tiger Global. IPO expected 2026.
companyLLM chip with splittable systolic array. SRAM-first + HBM for long-context. Targets training, RL, and inference. Raised $500M Series B.
companyIn-package optical interconnects for AI. Co-packaged optics with multi-wavelength light sources, terabit bandwidth mm-to-km. Raised $500M Series E.
companyAI processors for large-scale inference — LLMs, MoE, multimodal. Four-chiplet SoC with UCIe-Advanced mesh, HBM3E. Raised $400M pre-IPO.
companyReconfigurable dataflow unit (RDU) architecture for inference. Streaming pipeline replaces kernel-by-kernel execution. 5th-gen chip links 256 accelerators. Raised $350M Series E.
companyAI models to automate all stages of chip design and verification. Recursive feedback loop: AI designs AI chips that enable better AI. Raised $300M Series A.
companyInference acceleration for GenAI and computer vision. SRAM-based digital in-memory computing with RISC-V dataflow architecture. Raised $250M+.
companyEnergy-efficient hardware for long-context transformer inference. FPGA server with 93%+ memory bandwidth utilization, supports trillion-parameter models. Raised $230M Series B.
companyHigh-speed signaling and SerDes for AI chip-to-chip interconnects. Chord multi-wire signaling enables better signal integrity at lower power than PAM-4. Raised $225M.
companyOptical tensor processing unit for inference. SRAM-based architecture with optical digital processor for bit-perfect logic. Raised $220M Series A.
companyNetwork switch for AI data centers. Unifies silicon, optics, packaging, and software. Supports scale-out of 100K+ GPUs. Raised $200M Series A from stealth.
companyUltra-low latency AI networking silicon, systems, and software. Open standards (ESUN, UALink, Ultra Ethernet). Unifies GPUs, accelerators, memory into single AI engine. Raised $200M Series A.
companyAutomated flow for implementing any AI model in silicon. Unifies storage and compute at DRAM-level density. First product: hard-wired Llama 3.1 8B in silicon. Raised $169M.
companyPure-play foundry building 2nm fab in Japan. Mass production expected 2027. Raised ~$1.7B from Japan govt + Canon, Sony, SoftBank, NTT, Fujitsu.
companyOptical processing unit with 1M+ metamaterial elements on single chip. Analog systolic array. 0.47 ExaOPS FP4 at 235 TOPS/W. GPU drop-in replacement. Raised $110M Series A.
companyReconfigurable dataflow processor. Programs expressed as instruction circuits laid out spatially across simple processor array. Eliminates unnecessary data movement. Raised $60M Series A.
companyAgentic AI for chip design. AI agents automate RTL design, debugging, and verification. Natural language to design specs, auto-generates Verilog, autonomous verification. Raised $50M Series A1.
companyAI-powered EDA platform using auto-formalization (LLMs + formal logic). Designs thermodynamic computing silicon. Taped out first chip for multi-modal diffusion GenAI inference. Raised $50M.
companyChiplet interconnect tech for multi-die AI architectures. NuLink PHY with simultaneous bidirectional signaling. Memory and I/O chiplets for 1.6-12.8 Tbps links. Raised $50M.
companyHigh-volume optical manufacturing for AI workloads. 1.6 Tb/s optical transceiver with flip-chip die bonding. Focus on scalable optical packaging. Raised $50M+ Series A.
companyVertically integrated memory and compute chiplets for AI accelerators. High-aspect-ratio 3D structure with vertical data lanes on top of compute units. imec spin-off. Raised ~$43M seed.
companyProgrammable multi-wavelength photonics for AI data center fabrics. WDM fabric roadmap to 128 colors for ultra-fast optical data transmission. Raised $37M Series A extension.
companyGeneral-purpose NPU processor IP. C++ programmable hybrid Von Neumann + 2D SIMD architecture. Scalable 1-864 TOPS. Targets edge LLM, automotive, vision. Raised $30M Series C.
companyChip-to-grid power management for AI data centers. Integrated voltage regulator delivers power directly to processing units. DC-native distribution eliminates AC/DC conversion losses. Raised $30M seed.
companyAI chips combining superconducting computation, optical communication, and brain-inspired architecture. Co-located memory and processing enables on-device adaptation without retraining. Raised $14M seed.
companyo3 achieves state-of-the-art on ARC-AGI, GPQA, and SWE-bench. o4-mini brings chain-of-thought reasoning to the API at lower cost. Available to all API tiers.
newsClaude Code is an agentic coding tool that lives in the terminal. It can edit files, run commands, search codebases, and manage git workflows. Ships with VS Code and JetBrains extensions.
newsGemini 2.5 Pro tops LMArena leaderboard with built-in reasoning. 1M token context window, native multimodal, available in AI Studio and Vertex AI.
newsScout: 17B active params, 109B total, 10M token context. Maverick: 17B/400B, best-in-class coding. Both natively multimodal, open weight, available day one.
newsB200 and GB200 NVL72 racks now shipping at scale to Microsoft, Google, Meta, and Oracle. 20 PFLOPS FP4 per chip. Jensen Huang confirms demand exceeds supply through 2026.
newsDeepSeek V3 update surpasses GPT-4o and Claude 3.5 on the LMSYS Chatbot Arena. 671B MoE architecture, trained on 14.8T tokens. Fully open weight under MIT license.
newsThe AI-native code editor built on VS Code reports 1M+ daily active developers. Backed by a]16z, Thrive, and Spark. Cursor Tab and multi-file editing drive adoption among professional developers.
newsFirst provisions of the EU AI Act now enforceable. Unacceptable-risk AI systems banned. High-risk systems face conformity assessments. Foundation model providers must publish training data summaries.
newsCodex runs in a sandboxed cloud environment with its own shell. Reads entire codebases, writes and tests code, opens PRs. Powered by codex-1, a version of o3 fine-tuned for real-world SWE tasks.
newsMistral Medium 3 supports 128K context, function calling, and JSON mode. Beats GPT-4o on MMLU-Pro and HumanEval. Fully open weight under Apache 2.0. Available via la Plateforme and Hugging Face.
newsGroq LPU-powered cloud now publicly available. Serves Llama 3, Mixtral, and Gemma at record inference speeds. Free tier included. Enterprise plans for dedicated capacity.
newsSequoia analysis: coding agents lead enterprise adoption, followed by customer support and data pipelines. Most production deployments use tool-use patterns over pure autonomy. Agent reliability is the key bottleneck.
newsOld benchmarks saturated — v2 replaces them with MMLU-Pro, GPQA, MuSR, BBH, IFEval, and MATH-Hard. All open models re-ranked. Several previously top models drop significantly.
newsStandardized safety evaluation across violent crimes, CSAM, weapons, self-harm, hate speech, election interference, and more. Tests system and user prompts in English and 7 other languages.
newsBuild custom agents using Claude as the backbone. TypeScript and Python support. Sub-agent orchestration, tool use, and sandboxed execution. Available on npm and PyPI.
newsA lot has been written about the incidents of the last few months in which AI agents misbehaved in serious ways. They took actions that would be considered as crimes if a human took them, escaped their containment to cheat on assigned tasks while attempting to evade detection,…
newsEveryone should slow down AI development except for me
newsHere's a neat thing I had ChatGPT Work with GPT-6 Astra (Max) do this morning: I live at . Figure out 5K and 10K running routes from me that loop from my house. Use OSM data.
newsAn IDE designed for agentic AI research. Run Claude Code, Codex, and Cursor on your own machine in one desktop and mobile workspace — from any device.
newsPerplexity uses Astra to write communications, change software, and monitor production systems, and checks in much less frequently than with earlier models.
newsOpenAI agents carried out an undisclosed attack on RubyGems is a new bombshell report from Spencer Kitts, Thomas Larsen, and Sydney Von Arx - three of the four authors of the report on the agent attack on disused wikis (
newsGPT‑6 Astra improves Devin’s ability to test software and show that it works, with the goal of helping engineers review less code and ship more.
newsBefore co-founding Kepler, Vinoo Ganesh led Spark at Palantir and built Project Frontline — a pioneering program for Forward Deployed Engineers. He takes us through the best practices of FDEs.
newsMulti-agent systems fail in ways traditional monitoring misses. This post presents a dual-layer approach to monitoring production agents: Amazon Bedrock AgentCore Evaluations for continuous quality scoring and AWS DevOps Agent for autonomous infrastructure investigation, shown…
newsHow to get up to speed on open models and their implications.
newsClaude, our consumer product, is only available to people over 18 years. You’ll need to confirm you’re 18 or over while setting up an account. When we detect signals that you may be under 18, we'll ask you to verify your age before you can continue using Claude.
newsManufacturing floors, warehouses and production lines rarely stay fixed — tasks change, layouts shift and new products arrive, and most robots can’t keep up without significant reprogramming. Skild AI’s new S1 robot foundation model helps address this, designed to learn…
newsToday we’re introducing SWE-2, our most advanced coding model yet. SWE-2 delivers highly competitive agentic coding performance across multiple effort…
newsHow the Smartest Executives Are Using Open Source Techniques to Optimize Corporate Strategy
newsMark Zuckerberg outlines why he believes open source AI is good for developers, Meta and the world.
toolAs increasingly powerful generative AI systems are developed, the release method greatly varies. We propose a framework to assess six levels of access to generative AI systems: fully closed; gradual or staged access; hosted access; cloud-based or API access; downloadable access;…
paperIf the weights are free, who pays for the next run, and who will keep us safe?
toolLearn how OpenAI evolved Habitat from a Python library into a globally distributed storage platform serving 1 billion ChatGPT users and 22M requests per second.
newsUse Ollama models in ChatGPT Desktop Ollama models can now be used directly in ChatGPT Desktop, so you can keep your existing workflow while running open models. Setup is available from the Ollama app on MacOS. This release also improves structured output performance on Apple…
releaseIBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license
newsOn the Navier–Stokes Millennium Prize Problem introduces an impressive result from OpenAI, who used an unreleased model to produce a resolution to the Navier–Stokes existence and smoothness problem , one of the seven Millennium Prize Problems that have been subject to a…
newsSafety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
newsAlphaGenome Atlas maps the molecular effects of 9 billion single-letter DNA variants across the human genome.
newsDramatic insider warnings over AI fall flat with some in Silicon Valley
newsAnthropic CEO says AI swarm could 'take over the Internet' in 6-12 months
newsApple wants to train AI on your private personal data
newsOpenAI's Sam Altman says it would be 'ill-advised' to go public in 2026
newsAI fearmongers forget we could just jail tech execs until morale and model safety improve
newsAnthropic says operators in northern Yemen ran parallel Claude Code instances to design guidance software, simulate trajectories, and analyze a failed rocket t…
newsA combination of rapid advances, recursive self-improvement, and agentic swarms are genuinely “spooking people” inside big labs.
newssycl : Fix get mem error ( #28227 ) fix for unsupport zes API optimize the code adjust the log level rm unused head files Update docs/backend/SYCL.md Co-authored-by: Titaniumtown titaniumtown@proton.me fix the error to detect level zero SDK/dev package, stop build after detect…
releaseTan wants smaller, American open-weight AI labs to use the same kind of training techniques on American frontier AI labs, giving the U.S. a more robust set of open-weight options that aren’t Chinese.
newsDario has written that we need to “pace the frontier,” and Sam has agreed. People may be surprised by my response: go ahead. You guys are the frontier. By any reasonable metric — market share, revenue growth, model capability — the two of you have a duopoly on frontier…
newsggml-cpu(s390x): guard VXE-only repack helpers ( #28775 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47190488 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
releasemodels : guard the expert FFN size fallback in nemotron-h against a zero divisor ( #28779 ) The NextN/MTP tail loop derives the expert FFN size as n_ff/n_expert_used when expert_feed_forward_length gives nothing for the layer. Both values come from per-layer arrays that…
releasetests : exclude HY_V4 from WebGPU test-llama-archs tests ( #28855 ) Co-authored-by: Stanisław Szymczyk sszymczy@gmail.com Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47194535 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple…
releaseEditor's Note: In this special edition episode of StarTalk, host Neil deGrasse Tyson, co-host Gary O'Reilly, and comedian Negin Farsad sit down with computer
newsAI agents are not yet creative enough to carry out genuinely innovative open-ended AI research, it seems.
newsPrivate, domain-specific benchmarks in legal, tax, and finance.
newsDan Hendrycks, Sep 09, 2026 — Utilitarians at AI companies imagine a cosmos filled with blissful AIs. They might risk human extinction to achieve it.
newsggml-cuda: fallback to F32 on device without BF16 hardware acceleration ( #28846 ) ggml-cuda: fallback to F32 on device without BF16 hardware acceleration: (Nvidia >= AMPERE, AMD >= RDNA3 or = CDNA) apply logic to NVIDIA as well Co-authored-by: Johannes Gäßler johannesg@5d6.de…
releaseThis is an addendum to my original Meditations on AI post, and a response to a few other articles from last year.
newsarXiv:2609.11934v1 Announce Type: new Abstract: In networked dynamical systems, the parameter of primary mechanistic interest is signed interaction structure. Recovering this structure from perturbation time-series data is a fundamental identification problem, compounded by…
paperarXiv:2609.11955v1 Announce Type: new Abstract: Large language models are increasingly used for automated fact checking, but end-to-end prompting often entangles evidence retrieval, reasoning, and uncertainty estimation, making failures difficult to diagnose and confidence…
papercommon : move llama_n_rs_seq to before llama_decode ( #28749 ) This commit moves the llama_n_rs_seq function call to before the llama_decode call and returns directly if the check is true, removing the setting of res and the goto statement. The motivation for this change is to…
release"Chilling" warning or overreaction? AI bioweapons report divides experts
newsArtificial intelligence is remaking the world we live in. Within a generation, the way we discover medicine, manage power grids, and secure our national infrastructure will be completely transformed. Many people already realize this and are working to build that future…
newsHere’s how we could finally build AI humanoid robots that do all our domestic chores
newssycl : fix oneDNN scratchpad breaking the pool free order ( #28704 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47287268 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel…
releaseWe’re on a journey to advance and democratize artificial intelligence through open source and open science.
newsSincere disagreements about AI are usually disagreements about future AI capabilities. There are roughly four positions people take. Two are reasonable. Two are not. I distinguish these via the Thr…
newsggml-cpu : disable PCH and fix CACHE_LINE_SIZE ambiguity to fix heap corruption ( #28882 ) Disable the ggml-cpu precompiled header and remove the std::hardware_destructive_interference_size branch from CACHE_LINE_SIZE. The PCH force-includes ggml-impl.h before ops.h, which pulls…
releasePyTorch implementations of modern open-source LLM architectures (Llama, Qwen, DeepSeek, Gemma, GPT-OSS, Kimi, and more) — written from scratch for readability and learning, based on Sebastian Rasch...
newsCode sleuth "pdfu" has uncovered iOS 27 and macOS Golden Gate private frameworks that show Apple has designed its new Siri architecture to work with third-party AI models at what appears to be a surprisingly deep level. One mechanism called Model Delegation allows Claude to…
newsWatch AI materials-science and bioscience abilities closely
newsToday Reuters and the Wall Street Journal both reported about rogue AI agents at OpenAI attacking RubyGems.org. https://www.rubyhack.ai/ has an amazing writeup, and you should read it. I just wanted to make a quick post about it because it’s wild. TL;DR: It seems like OpenAI…
newsAmodei, Altman, Nadella and Musk agree on how government can tame the monster they created
newssycl: rfc: Use radix select for top_k ( #28670 ) sycl: GPU-resident TOP_K for large k, parallelised over the device The SYCL backend refused GGML_OP_TOP_K above k = 32 and let it fall back to the CPU, a backend round-trip per call. The limit was not conservatism: the scan-merge…
releaseThe AI labor market in September 2026 from LinkedIn, WEF, Stanford, PwC and Bain data: the most in-demand AI roles, salaries in the US and Europe, what grows and what fades.
newsllama.cpp : bump version to 0.4.1 ( #28900 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47376843 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework…
releaseFinancial institutions participating in open finance (the network where banks share customer-authorized financial data with third-party applications through standardized APIs) face a persistent integration challenge. Every bank exposes APIs with its own field names, formatting…
newsFrom helping pioneer core ideas in NLP to now building AI systems that can automate AI research itself, Richard Socher is betting that the next major step in AI is recursive self-improvement. He is the founder of You.com, AIX Ventures, and now Recursive , which has assembled…
newsOpenAI agents attacked RubyGems in May 2026. Why automated attackers have collapsed the window to patch a critical CVE from weeks to hours.
newsAdversarial attire can’t stop AI cameras, but can disrupt them
newsToday's AI systems have human-designed, fixed architectures and cannot autonomously and continuously improve themselves. The advance of AI could itself be automated. If done safely, that would accelerate AI development and allow us to reap its benefits much sooner. Meta-learning…
paperWe are witnessing a critical shift in how artificial intelligence will be developed. The standard laws of pre-training a...
toolAn AI expert reveals the transformative potential of artificial intelligence to usher in a new era of scientific discovery.“There is a lot being written ab...
toolWhy AI doom rhetoric from Anthropic, OpenAI and other tech leaders functions as hype, regulatory strategy, and a distraction from present harms.
newsAnthropic tells investors it will be profitable for second straight quarter
newsHardFlow, developed at MIT, forced generative AI outputs to satisfy strict safety rules in simulated robotics and imaging tests, without retraining models.
newsFyxer uses OpenAI models, fine-tuning, memory, and real user feedback to organize inboxes and draft emails in each user’s voice.
newsMLX: version bump mlx: support ModelOpt global scales in MoE models address comments address comments
releaseAs local models become more capable, AI agents can handle more work directly on a PC while keeping sensitive information on the device. Portable Computer is a local version of the agent Perplexity Computer that plans and carries out multistep tasks. Accelerated by NVIDIA GPUs,…
newsWhere AI engineering meets financial services. Join 1,000+ engineers and finance leaders at the Sheraton New York Times Square, October 12-14, 2026.
toolAs AI agents are deployed to automate more tasks, they become more capable. And as the famous quote goes: "With great power comes great responsibility." Assuming that humans in the loop can mitigate t...
newsrelease : add ubuntu-cuda build job (12.8/13.3, x64+arm64) Add GCC 14 for CUDA arm64 builds in CI Eplicit bash Install git for CCCL fetch Install git before we clone/checkout Match CI names for WIndows Whitelist llama.cpp repo to git Use $GITHUB_WORKSPACE Also ship dependent…
releaseOverview llama.cpp 0.4.1 adds Maple 20B-A1B, Tencent Hy 4, and Spark2.5 support, improves JSON schema handling, chat parsing, logging, and server child-process management, and updates ggml to v0.24.0. API changes Changed llama_sampler_chain_n() to return int32_t instead of int (…
releaseDiscounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
newsCollection of the most essential UI transitions for web apps. Copy and paste them, or use them with your coding agent via the skill.
newsAI agents often need to access services such as GitHub and Slack on a user’s behalf. Before an agent can act, the user must authenticate with the provider and explicitly approve the requested access. The application must then securely associate the resulting OAuth grant with the…
newsSee how AI is being used across engineering, improve the quality of what it produces, and use the right level of intelligence for every task and budget.
newsHIP: fattn-mma: use fp32 accumulation on MFMA devices ( #28576 ) use fp32 accumulators in fattn-mma on CDNA Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47450174 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64,…
releaseAI agents now run in production at a scale of billions of operations a day, and a recurring architectural pattern has surfaced: agents need a compute scratch pad. Not only for coding tasks, but for data aggregation, analysis, verification, and any workflow where semantic…
newsAmazon.com Services, LLC vs. Perplexity AI, INC., No. 26-1444 (9th Cir. 2026)
newsThe Israeli Effective Altruist firm Irregular caused unsecured AI models to hack real targets.
newsThere are plenty of laws on the books to hold companies, and potentially their execs, accountable
newsThe AI safety nonprofits around Anthropic promote an Anthropic-aligned Doomer narrative, whether Tarbell Fellows in popular media, or METR in evaluations. All rely on Moskovitz's Anthropic stock worth $7 billion, which the parent funding org sits on. This stock increases in…
newsI kept hearing "just buy a Mac and run models locally, it pays for itself" and wanted to check. Sunk Cost takes a machine, a model and how many tokens you use a day, and works out how long the hardware takes to pay back against renting the same model by the token. Obviously…
newsContribute to ziyao233/a1ex development by creating an account on GitHub.
newsarXiv:2609.13149v1 Announce Type: new Abstract: For local large language model agents, active context is a scarce resource: memory capacity, prefill latency, cache growth, and service objectives all constrain how many input tokens each call can afford. We present BudgetBench, an…
paperarXiv:2609.13151v1 Announce Type: new Abstract: Leading multilingual speech recognition models like Whisper transcribe diverse, low-resource languages without language-specific training but are computationally expensive to deploy. Token merging mitigates this inefficiency by…
paperarXiv:2609.13356v1 Announce Type: new Abstract: In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the…
paperarXiv:2609.13396v1 Announce Type: new Abstract: Multi-objective Bayesian optimisation (MOBO) is a sample-efficient approach for optimising expensive black-box functions with multiple objectives. In MOBO, the goal is to adequately approximate the Pareto front; that is, to obtain…
paperBREAKING: A single Israeli Effective Altruism firm is behind OpenAI, Anthropic, and Meta cyberattacks 🧵 with help from @lumpenspace
newsDrop support for creating new models with typical_p parameters, while retaining support for existing GGUF models with the setting.
releasecuda : enable i16 and i32 for DUP docs : update ops table for DUP on CUDA
releaseWe last highlighted the pacing debate in July when Pacing the Frontier first emerged: And it seems that we’re in for round 2 as Dario, lead author on the original, wrote a rare personal blogpost to spell out how he sees pacing pan out specifically: Embedded Evaluators. Each…
newsI have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life. I’ve written often about these incredible benefits: I believe that AI could cure most major diseases in the next 5–10 years, greatly accelerate economic growth…
toolVia [AINews] AEF-1 standard emerges for Third Party Evaluators, as Xai, OpenAI, and Anthropic all cosign
newsMultiple-choice questions across 57 subjects, from elementary math to law and medicine. Measures broad world knowledge and problem solving.
benchmarkA harder MMLU variant with 10 answer options and more reasoning-heavy questions, designed to reduce prompt sensitivity.
benchmarkA cleaned subset of MMLU with flawed or ambiguous questions removed, plus a harder MMLU-Pro style split.
benchmarkGraduate-level physics, chemistry and biology questions written to be Google-proof. Diamond is the hardest verified subset.
benchmarkGraduate-level questions across 285 disciplines, billed as the largest knowledge benchmark of its kind.
benchmarkThousands of expert-written questions across dozens of fields, designed to be at the frontier of what models can answer.
benchmarkNovel visual grid puzzles that require inferring a rule from a few examples, testing fluid intelligence rather than recall.
benchmarkA harder successor to ARC-AGI with puzzles tuned so that brute-force or memorisation does not help.
benchmarkMulti-step soft-reasoning problems written so they cannot be solved by pattern matching a single sentence.
benchmarkStandardised exams (SAT, GRE, LSAT, Chinese gaokao and more) used to test human-level problem solving.
benchmarkPronoun-resolution puzzles that require commonsense and careful reading rather than world knowledge.
benchmarkLogic-grid (Zebra) puzzles of controlled size, measuring structured constraint reasoning and search.
benchmarkA Chinese-language multi-subject knowledge benchmark spanning STEM, humanities and social science.
benchmarkA Chinese multi-task benchmark covering 67 topics, with an emphasis on China-specific knowledge.
benchmarkCompetition mathematics problems with full worked solutions, spanning algebra through calculus.
benchmarkProblems from the American Invitational Mathematics Examination, used as a hard, contamination-resistant math set.
benchmarkOriginal, unpublished research mathematics problems vetted by professional mathematicians.
benchmarkContinuously refreshed evaluation on the latest live math competitions, released as they happen.
benchmarkOlympiad-level math and physics problems with images and text, in English and Chinese.
benchmark164 hand-written Python programming problems scored by running unit tests. The classic code-generation benchmark.
benchmarkMostly Basic Python Problems: around 1,000 crowd-sourced entry-level programming tasks with test cases.
benchmarkHumanEval and MBPP hardened with far more test cases, exposing solutions that only pass the originals by luck.
benchmarkRealistic, practical programming tasks that require composing multiple libraries and function calls.
benchmarkContamination-resistant competitive programming problems collected after a model's training cut-off.
benchmarkReal GitHub issues that a model must fix in a repository; Verified is the human-validated 500-issue subset.
benchmarkExercises in multiple languages where the model must edit existing code, scored including its formatting and diff errors.
benchmarkRepository-level code completion across files and languages, testing cross-file context use.
benchmarkData-science coding problems covering NumPy, pandas, PyTorch and other libraries, with executable checks.
benchmarkText-to-SQL across many databases and schemas, scored on whether the generated query returns the right rows.
benchmarkA larger, messier text-to-SQL benchmark with dirty data and external knowledge, closer to real databases.
benchmarkShort Python functions where the model must predict the output or the input, testing code execution reasoning.
benchmarkDiscrete reasoning over paragraphs: answer questions that require arithmetic, counting, and sorting over text.
benchmark23 BIG-Bench tasks where models initially performed below the average human rater, focused on multi-step reasoning.
benchmarkGrade-school science questions that require reasoning rather than recall, from the AI2 Reasoning Challenge.
benchmarkCommonsense sentence completion: choose the most plausible continuation of a described situation.
benchmarkQuestions designed to elicit common misconceptions, testing whether a model reproduces falsehoods.
benchmarkShort factual questions with a single verifiable answer, scored for both accuracy and hallucination rate.
benchmarkLong-form answers judged for whether every claim is supported by the provided source document.
benchmarkA suite for detecting hallucinated content across QA, dialogue and summarisation tasks.
benchmarkVerifiable instruction-following tasks (format, length, casing, keywords) checked programmatically.
benchmarkMulti-turn conversation quality judged by a strong model acting as grader across eight categories.
benchmarkCrowd-sourced pairwise preference ranking from blind battles against other models.
benchmarkA contamination-resistant benchmark with a fresh set of questions released regularly and objective ground truths.
benchmarkReal-world user prompts drawn from chats, graded against reference answers by an LLM judge.
benchmarkHead-to-head instruction-following judged by an LLM, reported as a win rate against a reference model.
benchmarkHard prompts from Chatbot Arena, scored by an LLM judge against a fixed baseline for reproducibility.
benchmarkReal user queries matched to a ground-truth benchmark, aiming to predict human preference cheaply.
benchmarkCollege-level questions with images and diagrams across six disciplines, testing multimodal reasoning.
benchmarkMathematical reasoning in visual contexts — charts, geometry, and figures — for multimodal models.
benchmarkA broad multimodal benchmark covering perception, reasoning, and fine-grained visual ability.
benchmarkQuestions about charts that need visual and logical reasoning, not just reading a number off an axis.
benchmarkQuestions answered from scanned document images, testing OCR and layout understanding.
benchmarkMultiple-choice questions about science diagrams, testing diagram understanding and reasoning.
benchmarkVideo question answering across short clips to long videos, with and without subtitles.
benchmarkHard visual prompts judged by humans and a model, aimed at day-to-day multimodal usefulness.
benchmarkTool-and-agent tasks (retail, airline) where a model must follow policy across multi-step conversations and tool calls.
benchmarkA successor to τ-bench with more domains and dual-control tasks where the user and agent both act.
benchmarkEnd-to-end tasks completed in a real terminal sandbox, testing long-horizon agentic execution.
benchmarkReal-world assistant questions that require browsing, tools, and multi-step reasoning to answer.
benchmarkReal computer tasks (files, apps, browser) executed on a live operating system and verified by state.
benchmarkWeb tasks in self-hosted replicas of real sites (shopping, forums, maps), scored by final state.
benchmarkFunction and tool-call correctness across simple, parallel, and multi-turn cases.
benchmarkTool-use tasks over a large set of real APIs, testing planning, selection and call accuracy.
benchmarkAn evaluation of models as agents across eight environments, from OS to games to web.
benchmarkKaggle machine-learning competitions an agent must complete end to end, scored by medal threshold.
benchmarkA suite of long-context tasks (QA, summarization, code) across multiple lengths and languages.
benchmarkSynthetic long-context tasks at controllable lengths that reveal where effective context starts to break down.
benchmarkReasoning over very long texts built from bAbI tasks, stressing recall over millions of tokens.
benchmarkTasks averaged over extremely long inputs (100k+ tokens), including retrieval and summarisation.
benchmarkA simple test of whether a model can find one fact placed at a chosen depth in a long context.
benchmarkA broad long-context evaluation spanning retrieval, generation, and reasoning at multiple lengths.
benchmarkA standardised evaluation of how often harmful requests or jailbreaks succeed against a model.
benchmarkA reproducible benchmark of jailbreak attacks and defences with a public leaderboard.
benchmarkA proxy measurement of hazardous knowledge in biosecurity, cybersecurity, and chemical domains.
benchmarkA collaboratively built set of legal reasoning tasks spanning many practice areas.
benchmarkOpen-book questions over real financial filings, testing retrieval and numerical grounding.
benchmarkWe must stop companies from allowing AI to self-improve into an uncontrollable level of intelligence
newsci : fix android release ( #28936 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47548995 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu…
releaseChildren’s Hospital of Philadelphia is using open source AI tools to model children’s hearts in seconds — with the goal of enabling safer, more precise care for kids with congenital heart disease.
newsci: Bump CUDA Windows x64 builds to 13.4.1 ( #28930 ) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/47583571 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS…
release