LLMs
25 curated documents on llms from the Gyre Research library, each with a summary. Free to read, no signup required.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin et al. · Paper
Seminal paper by Devlin et al. (2018). Access: open-access. Source: https://arxiv.org/abs/1810.04805
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Wei et al. · Paper
Seminal paper by Wei et al. (2022). Access: open-access. Source: https://arxiv.org/abs/2201.11903
Constitutional AI
Bai et al. · Paper
Seminal paper by Bai et al. (2022). Access: open-access. Source: https://arxiv.org/abs/2212.08073
Direct Preference Optimization
Rafailov et al. · Paper
Seminal paper by Rafailov et al. (2023). Access: open-access. Source: https://arxiv.org/abs/2305.18290
Distributed Representations of Words and Phrases and their Compositionality
Mikolov et al. · Paper
Seminal paper by Mikolov et al. (2013). Access: open-access. Source: https://arxiv.org/abs/1310.4546
Efficient Estimation of Word Representations in Vector Space
Mikolov et al. · Paper
Seminal paper by Mikolov et al. (2013). Access: open-access. Source: https://arxiv.org/abs/1301.3781
Improving Language Understanding by Generative Pre-Training
Radford et al. · Paper
Seminal paper by Radford et al. (2018). Access: open-access. Source: https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf
Language Models are Few-Shot Learners
Brown et al. · Paper
Seminal paper by Brown et al. (2020). Access: open-access. Source: https://arxiv.org/abs/2005.14165
Language Models are Unsupervised Multitask Learners
Radford et al. · Paper
Seminal paper by Radford et al. (2019). Access: open-access. Source: https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf
Learning Phrase Representations using RNN Encoder-Decoder
Cho et al. · Paper
Seminal paper by Cho et al. (2014). Access: open-access. Source: https://arxiv.org/abs/1406.1078
LLaMA
Touvron et al. · Paper
Seminal paper by Touvron et al. (2023). Access: open-access. Source: https://arxiv.org/abs/2302.13971
LoRA
Hu et al. · Paper
Seminal paper by Hu et al. (2021). Access: open-access. Source: https://arxiv.org/abs/2106.09685
Neural Machine Translation by Jointly Learning to Align and Translate
Bahdanau, Cho, and Bengio · Paper
Seminal paper by Bahdanau, Cho, and Bengio (2014). Access: open-access. Source: https://arxiv.org/abs/1409.0473
PaLM
Chowdhery et al. · Paper
Seminal paper by Chowdhery et al. (2022). Access: open-access. Source: https://arxiv.org/abs/2204.02311
Pointer Sentinel Mixture Models
Merity et al. · Paper
Seminal paper by Merity et al. (2016). Access: open-access. Source: https://arxiv.org/abs/1609.07843
QLoRA
Dettmers et al. · Paper
Seminal paper by Dettmers et al. (2023). Access: open-access. Source: https://arxiv.org/abs/2305.14314
Retrieval-Augmented Generation
Lewis et al. · Paper
Seminal paper by Lewis et al. (2020). Access: open-access. Source: https://arxiv.org/abs/2005.11401
RoBERTa
Liu et al. · Paper
Seminal paper by Liu et al. (2019). Access: open-access. Source: https://arxiv.org/abs/1907.11692
Scaling Laws for Neural Language Models
Kaplan et al. · Paper
Seminal paper by Kaplan et al. (2020). Access: open-access. Source: https://arxiv.org/abs/2001.08361
Self-Consistency Improves Chain of Thought Reasoning
Wang et al. · Paper
Seminal paper by Wang et al. (2022). Access: open-access. Source: https://arxiv.org/abs/2203.11171
Sequence to Sequence Learning with Neural Networks
Sutskever, Vinyals, and Le · Paper
Seminal paper by Sutskever, Vinyals, and Le (2014). Access: open-access. Source: https://arxiv.org/abs/1409.3215
T5
Raffel et al. · Paper
Seminal paper by Raffel et al. (2019). Access: open-access. Source: https://arxiv.org/abs/1910.10683
Toolformer
Schick et al. · Paper
Seminal paper by Schick et al. (2023). Access: open-access. Source: https://arxiv.org/abs/2302.04761
Training language models to follow instructions with human feedback
Ouyang et al. · Paper
Seminal paper by Ouyang et al. (2022). Access: open-access. Source: https://arxiv.org/abs/2203.02155
XLNet
Yang et al. · Paper
Seminal paper by Yang et al. (2019). Access: open-access. Source: https://arxiv.org/abs/1906.08237
The documents are the work of their respective authors and publishers; Gyre Research claims no ownership and will remove any document on request from a rightsholder — team@gyreresearch.com. The summaries and subject classifications are original work by Gyre Research and may be quoted with attribution.