Attention Is All You Need
The foundational paper introducing the Transformer architecture that powers all modern LLMs
paperCore papers and the open-weight releases they led to.
The foundational paper introducing the Transformer architecture that powers all modern LLMs
paperBidirectional encoder representations that revolutionized NLP transfer learning
paperThe GPT-3 paper showing that scaling language models enables few-shot learning across tasks
paperParameter-efficient fine-tuning that reduces trainable parameters by orders of magnitude
paperPrompting LLMs to think step-by-step dramatically improves reasoning performance
paperTraining helpful, harmless, and honest AI systems using AI-generated feedback
paperMeta's open-source LLM family from 7B to 70B parameters, available for research and commercial use
modelCompact 7B model outperforming Llama 2 13B with grouped-query and sliding window attention
model