References
Agarwal, Rishabh, Nino Vieillard, Yongchao Zhou, et al. 2023.
“On-Policy Distillation of Language Models: Learning from
Self-Generated Mistakes.” arXiv Preprint
arXiv:2306.13649.
Aggarwal, Pranjal, and Sean Welleck. 2025. “L1: Controlling How
Long a Reasoning Model Thinks with Reinforcement Learning.”
arXiv Preprint arXiv:2503.04697.
Agrawal, Amey, Nitin Kedia, Ashish Panwar, et al. 2024. “Taming
Throughput-Latency Tradeoff in LLM Inference with
Sarathi-Serve.” 18th USENIX Symposium on Operating Systems
Design and Implementation (OSDI).
Agrawal, Lakshya A. et al. 2025.
“GEPA: Reflective Prompt Evolution Can Outperform
Reinforcement Learning.” arXiv Preprint
arXiv:2507.19457.
Ahia, Orevaoghene, Sachin Kumar, Hila Gonen, et al. 2023. “Do All
Languages Cost the Same? Tokenization in the Era of Commercial Language
Models.” Proceedings of the 2023 Conference on Empirical
Methods in Natural Language Processing.
Ahn, Michael, Anthony Brohan, Noah Brown, et
al. 2022. “Do as i Can, Not as i Say: Grounding Language in
Robotic Affordances.” arXiv Preprint arXiv:2204.01691.
Ainslie, Joshua, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy,
Federico Lebrón, and Sumit Sanghai. 2023. “GQA: Training
Generalized Multi-Query Transformer Models from Multi-Head
Checkpoints.” Proceedings of the Conference on Empirical
Methods in Natural Language Processing.
Amershi, Saleema, Dan Weld, Mihaela Vorvoreanu, et
al. 2019. “Guidelines for Human-AI Interaction.”
Proceedings of the CHI Conference on Human Factors in Computing
Systems.
Anthropic. 2024a. Building Effective Agents. Https://www.anthropic.com/research/building-effective-agents.
Anthropic. 2024b. Introducing the Model Context Protocol. Https://modelcontextprotocol.io.
Arditi, Andy et al. 2024. “Refusal in
Language Models Is Mediated by a Single Direction.” arXiv
Preprint arXiv:2406.11717.
Armeni, Iro, Zhi-Yang He, JunYoung Gwak, et al. 2019.
“3D Scene Graph: A Structure for Unified Semantics,
3D Space, and Camera.” IEEE International
Conference on Computer Vision (ICCV).
Asai, Akari, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh
Hajishirzi. 2024. “Self-RAG: Learning to Retrieve, Generate, and
Critique Through Self-Reflection.” International Conference
on Learning Representations.
Ashkboos, Saleh, Maximilian L. Croci, Marcelo Gennari do Nascimento,
Torsten Hoefler, and James Hensman. 2024. “SliceGPT:
Compress Large Language Models by Deleting Rows and Columns.”
International Conference on Learning Representations.
Association for Computing Machinery. 2018. ACM Code of Ethics and
Professional Conduct. Https://www.acm.org/code-of-ethics.
Azar, Mohammad Gheshlaghi, Zhaohan Daniel Guo, Bilal Piot, et al. 2024.
A General Theoretical Paradigm to Understand Learning from Human
Preferences.
Bai, Yuntao, Saurav Kadavath, Sandipan Kundu, et
al. 2022. “Constitutional AI: Harmlessness
from AI Feedback.” arXiv Preprint
arXiv:2212.08073.
Barnett, Scott, Stefanus Kurniawan, Srikanth Thudumu, Zach Brannelly,
and Mohamed Abdelrazek. 2024. Seven Failure Points When Engineering
a Retrieval Augmented Generation System. https://arxiv.org/abs/2401.05856.
Basiri, Ali, Niosha Behnam, Ruud de Rooij, et al. 2016. “Chaos
Engineering.” In IEEE Software, No. 3, vol. 33. Https://principlesofchaos.org/.
Belrose, Nora et al. 2023. “Eliciting
Latent Predictions from Transformers with the Tuned Lens.”
arXiv Preprint arXiv:2303.08112.
Besta, Maciej, Nils Blach, Ales Kubicek, Robert
Gerstenberger, et al. 2024. “Graph of Thoughts: Solving
Elaborate Problems with Large Language Models.” Proceedings
of the AAAI Conference on Artificial Intelligence.
Betley, Jan et al. 2025. “Emergent
Misalignment: Narrow Finetuning Can Produce Broadly Misaligned
LLMs.” arXiv Preprint arXiv:2502.17424.
Beyer, Betsy, Chris Jones, Jennifer Petoff, and Niall Richard Murphy.
2016. Site Reliability Engineering: How Google Runs Production
Systems. Https://sre.google/sre-book/table-of-contents/; O’Reilly
Media.
Beyer, Betsy, Niall Richard Murphy, David K. Rensin, Kent Kawahara, and
Stephen Thorne. 2018. The Site Reliability Workbook: Practical Ways
to Implement SRE. Https://sre.google/workbook/table-of-contents/; O’Reilly
Media.
Black, Kevin, Noah Brown, Danny Driess, et
al. 2024. “π0: A
Vision-Language-Action Flow Model for General Robot Control.”
arXiv Preprint arXiv:2410.24164.
Bourtoule, Lucas, Varun Chandrasekaran, Christopher A. Choquette-Choo,
et al. 2021. “Machine Unlearning.” IEEE Symposium on
Security and Privacy.
Bradley, Ralph Allan, and Milton E. Terry. 1952. “Rank Analysis of
Incomplete Block Designs: I. The Method of Paired Comparisons.”
Biometrika 39 (3/4): 324–45.
Brohan, Anthony, Noah Brown, Justice Carbajal, et
al. 2022. “RT-1: Robotics Transformer for
Real-World Control at Scale.” arXiv Preprint
arXiv:2212.06817.
Brohan, Anthony, Noah Brown, Justice Carbajal, et
al. 2023. “RT-2: Vision-Language-Action Models
Transfer Web Knowledge to Robotic Control.” arXiv Preprint
arXiv:2307.15818.
Brown, Bradley, Jordan Juravsky, Ryan Ehrlich, et al. 2024. “Large
Language Monkeys: Scaling Inference Compute with Repeated
Sampling.” arXiv Preprint arXiv:2407.21787.
Brown, Tom B., Benjamin Mann, Nick Ryder, et
al. 2020. “Language Models Are Few-Shot Learners.”
Advances in Neural Information Processing Systems (NeurIPS).
Buçinca, Zana, Maja Barbara Malaya, and Krzysztof Z. Gajos. 2021.
“To Trust or to Think: Cognitive Forcing Functions Can Reduce
Overreliance on AI in AI-Assisted Decision-Making.”
Proceedings of the ACM on Human-Computer Interaction (CSCW).
Buhl, Marie Davidsen et al. 2024.
“Safety Cases for Frontier AI.” arXiv Preprint
arXiv:2410.21572.
Carlini, Nicholas, Daniel Paleka, Krishnamurthy Dj
Dvijotham, et al. 2024. “Stealing Part of a Production
Language Model.” International Conference on Machine
Learning.
Carlini, Nicholas, Florian Tramer, Eric Wallace, et al. 2021.
“Extracting Training Data from Large Language Models.”
USENIX Security Symposium.
Cemri, Mert, Melissa Z. Pan, Shuyi Yang, et al. 2025. “Why Do
Multi-Agent LLM Systems Fail?” arXiv Preprint
arXiv:2503.13657.
Chan, Brian J., Chao-Ting Chen, Jui-Hung Cheng, and Hen-Hsen Huang.
2024. “Don’t Do RAG: When Cache-Augmented Generation Is All You
Need for Knowledge Tasks.” arXiv Preprint
arXiv:2412.15605.
Chen, Charlie, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste
Lespiau, Laurent Sifre, and John Jumper. 2023. “Accelerating Large
Language Model Decoding with Speculative Sampling.” arXiv
Preprint arXiv:2302.01318.
Chen, Mark, Jerry Tworek, Heewoo Jun, et al.
2021. “Evaluating Large Language Models Trained on Code.”
arXiv Preprint arXiv:2107.03374.
Chen, Shouyuan, Sherman Wong, Liangjian Chen, and Yuandong Tian. 2023.
“Extending Context Window of Large Language Models via Positional
Interpolation.” arXiv Preprint arXiv:2306.15595.
Chhikara, Prateek et al. 2025.
“Mem0: Building Production-Ready AI
Agents with Scalable Long-Term Memory.” arXiv Preprint
arXiv:2504.19413.
Chowdhery, Aakanksha, Sharan Narang, Jacob Devlin,
et al. 2023. “PaLM: Scaling Language Modeling with
Pathways.” Journal of Machine Learning Research 24
(240): 1–113.
Chroma Research. 2025. Context Rot: How Increasing Input Tokens
Impacts LLM Performance. Chroma. https://www.trychroma.com/research/context-rot.
Clymer, Joshua et al. 2024. “Safety
Cases: How to Justify the Safety of Advanced AI Systems.”
arXiv Preprint arXiv:2403.10462.
Coalition for Content Provenance and Authenticity. 2024. C2PA
Technical Specification, Version 2.4. Https://spec.c2pa.org/specifications/specifications/2.4/specs/C2PA_Specification.html.
Cohen, Jacob. 1960. “A Coefficient of Agreement for Nominal
Scales.” Educational and Psychological Measurement 20
(1): 37–46.
Cormack, Gordon V., Charles L. A. Clarke, and Stefan Buettcher. 2009.
“Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank
Learning Methods.” Proceedings of the 32nd International ACM
SIGIR Conference on Research and Development in Information
Retrieval, 758–59.
Cui, Ganqu, Yuchen Zhang, Jiacheng Chen, et
al. 2025. “The Entropy Mechanism of Reinforcement Learning
for Reasoning Language Models.” arXiv Preprint
arXiv:2505.22617.
Dao, Tri, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.
2022. “FlashAttention: Fast and Memory-Efficient Exact Attention
with IO-Awareness.” Advances in Neural Information Processing
Systems 35.
Dao, Tri, and Albert Gu. 2024. “Transformers Are
SSMs: Generalized Models and Efficient Algorithms Through
Structured State Space Duality.” International Conference on
Machine Learning (ICML).
Dean, Jeffrey, and Luiz André Barroso. 2013. “The Tail at
Scale.” Communications of the ACM 56 (2): 74–80.
Debenedetti, Edoardo, Ilia Shumailov, Tianqi Fan, et al. 2025.
“Defeating Prompt Injections by Design.” arXiv Preprint
arXiv:2503.18813.
DeepSeek-AI. 2024a. “DeepSeek-V2: A Strong, Economical, and
Efficient Mixture-of-Experts Language Model.” arXiv Preprint
arXiv:2405.04434.
DeepSeek-AI. 2024b. “DeepSeek-V3 Technical
Report.” arXiv Preprint arXiv:2412.19437.
DeepSeek-AI. 2025. “DeepSeek-V3.2: Boosting
Long-Context Efficiency with DeepSeek Sparse
Attention.” arXiv Preprint arXiv:2512.02556.
DeepSeek-AI, Aixin Liu, Bei Feng, et al.
2024. “DeepSeek-V3 Technical Report.” arXiv Preprint
arXiv:2412.19437.
Deerwester, Scott, Susan T. Dumais, George W. Furnas, Thomas K.
Landauer, and Richard Harshman. 1990. “Indexing by Latent Semantic
Analysis.” Journal of the American Society for Information
Science 41 (6): 391–407.
Défossez, Alexandre, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2023.
“High Fidelity Neural Audio Compression.” Transactions
on Machine Learning Research. https://arxiv.org/abs/2210.13438.
Défossez, Alexandre, Laurent Mazaré, Manu Orsini, Amélie Rózière, and
Kyutai Team. 2024. Moshi: A Speech-Text Foundation Model for
Real-Time Dialogue. https://arxiv.org/abs/2410.00037.
Dettmers, Tim, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer.
2023. “QLoRA: Efficient Finetuning of Quantized
LLMs.” Advances in Neural Information Processing
Systems.
Dong, Yixin, Charlie F. Ruan, Yaxing Cai, et al. 2024. XGrammar:
Flexible and Efficient Structured Generation Engine for Large Language
Models. https://arxiv.org/abs/2411.15100.
Douillard, Arthur, Qixuan Feng, Andrei A. Rusu, et al. 2023.
“DiLoCo: Distributed Low-Communication Training of Language
Models.” arXiv Preprint arXiv:2311.08105.
Du, Mingxuan et al. 2025.
“DeepResearch Bench: A Comprehensive Benchmark for
Deep Research Agents.” arXiv Preprint arXiv:2506.11763.
Du, Yilun, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor
Mordatch. 2023. “Improving Factuality and Reasoning in Language
Models Through Multiagent Debate.” arXiv Preprint
arXiv:2305.14325.
Dwork, Cynthia, and Aaron Roth. 2014. “The Algorithmic Foundations
of Differential Privacy.” Foundations and Trends in
Theoretical Computer Science 9 (3–4).
Edge, Darren, Ha Trinh, Newman Cheng, et al. 2024. “From Local to
Global: A Graph RAG Approach to Query-Focused Summarization.”
arXiv Preprint arXiv:2404.16130.
Elhage, Nelson et al. 2021. “A
Mathematical Framework for Transformer Circuits.” Transformer
Circuits Thread.
Elhage, Nelson et al. 2022. “Toy
Models of Superposition.” Transformer Circuits Thread.
Es, Shahul, Jithin James, Luis Espinosa-Anke, and Steven Schockaert.
2023. RAGAS: Automated Evaluation of Retrieval Augmented
Generation. https://arxiv.org/abs/2309.15217.
Ethayarajh, Kawin, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and
Douwe Kiela. 2024. “KTO: Model Alignment as Prospect
Theoretic Optimization.” arXiv Preprint
arXiv:2402.01306.
Farquhar, Sebastian, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. 2024.
“Detecting Hallucinations in Large Language Models Using Semantic
Entropy.” Nature.
Faysse, Manuel, Hugues Sibille, Tony Wu, et al. 2025. “ColPali:
Efficient Document Retrieval with Vision Language Models.”
International Conference on Learning Representations.
Federal Trade Commission. 2024. Approaches to Address
AI-Enabled Voice Cloning. Https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/04/approaches-address-ai-enabled-voice-cloning.
Fedus, William, Barret Zoph, and Noam Shazeer. 2022. “Switch
Transformers: Scaling to Trillion Parameter Models with Simple and
Efficient Sparsity.” Journal of Machine Learning
Research 23 (120): 1–39.
Feng, Lang et al. 2025.
“Group-in-Group Policy Optimization for LLM Agent
Training.” arXiv Preprint arXiv:2505.10978.
Frantar, Elias, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2023.
“GPTQ: Accurate Post-Training Quantization for
Generative Pre-Trained Transformers.” International
Conference on Learning Representations.
Garcia-Molina, Hector, and Kenneth Salem. 1987. “Sagas.”
ACM SIGMOD Record 16 (3): 249–59.
Geng, Saibo, Hudson Cooper, Michael Moss, Cameron
Ross, Luca Beurer-Kellner, et al. 2025. Generating Structured
Outputs from Language Models: Benchmark and Studies. https://arxiv.org/abs/2501.10868.
Gerstgrasser, Matthias, Rylan Schaeffer, Apratim
Dey, et al. 2024. “Is Model Collapse Inevitable? Breaking
the Curse of Recursion by Accumulating Real and Synthetic Data.”
arXiv Preprint arXiv:2404.01413.
Gloeckle, Fabian, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz,
and Gabriel Synnaeve. 2024. “Better & Faster Large Language
Models via Multi-Token Prediction.” arXiv Preprint
arXiv:2404.19737.
Google. 2025. Agent2Agent (A2A) Protocol Specification. Https://a2a-protocol.org.
Google DeepMind. 2024. SynthID: Identifying
AI-Generated Content. Https://deepmind.google/models/synthid/.
Grattafiori, Aaron, Abhimanyu Dubey, Abhinav
Jauhri, et al. 2024. “The Llama 3 Herd of Models.”
arXiv Preprint arXiv:2407.21783.
Greenblatt, Ryan et al. 2024.
“Alignment Faking in Large Language Models.” arXiv
Preprint arXiv:2412.14093.
Greshake, Kai, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres,
Thorsten Holz, and Mario Fritz. 2023. “Not What You’ve Signed up
for: Compromising Real-World LLM-Integrated Applications with Indirect
Prompt Injection.” Proceedings of the 16th ACM Workshop on
Artificial Intelligence and Security.
Gu, Albert, and Tri Dao. 2023. “Mamba: Linear-Time Sequence
Modeling with Selective State Spaces.” arXiv Preprint
arXiv:2312.00752.
Gu, Albert, Karan Goel, and Christopher Ré. 2022. “Efficiently
Modeling Long Sequences with Structured State Spaces.”
International Conference on Learning Representations (ICLR).
Gu, Yuxian, Li Dong, Furu Wei, and Minlie Huang. 2023.
“MiniLLM: Knowledge Distillation of Large Language
Models.” arXiv Preprint arXiv:2306.08543.
Guo, Daya, Dejian Yang, Haowei Zhang, Junxiao Song,
et al. 2025. “DeepSeek-R1: Incentivizing
Reasoning Capability in LLMs via Reinforcement
Learning.” arXiv Preprint arXiv:2501.12948.
Haan, Pim de, Dinesh Jayaraman, and Sergey Levine. 2019. “Causal
Confusion in Imitation Learning.” Advances in Neural
Information Processing Systems (NeurIPS).
Hao, Shibo, Yi Gu, Haodi Ma, et al. 2023. “Reasoning with Language
Model Is Planning with World Model.” Conference on Empirical
Methods in Natural Language Processing.
Hardy, Norm. 1988. “The Confused Deputy: (Or Why Capabilities
Might Have Been Invented).” ACM SIGOPS Operating Systems
Review 22 (4): 36–38.
Hewitt, John, and Percy Liang. 2019. “Designing and Interpreting
Probes with Control Tasks.” Empirical Methods in Natural
Language Processing.
Hines, Keegan, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan
Zunger, and Emre Kiciman. 2024. “Defending Against Indirect Prompt
Injection Attacks with Spotlighting.” arXiv Preprint
arXiv:2403.14720.
Ho, Jonathan, Ajay Jain, and Pieter Abbeel. 2020. “Denoising
Diffusion Probabilistic Models.” Advances in Neural
Information Processing Systems (NeurIPS). https://arxiv.org/abs/2006.11239.
Hoffmann, Jordan et al. 2022.
“Training Compute-Optimal Large Language Models.”
Advances in Neural Information Processing Systems.
Holtzman, Ari, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020.
“The Curious Case of Neural Text Degeneration.”
International Conference on Learning Representations.
Hong, Jiwoo, Noah Lee, and James Thorne. 2024. “ORPO:
Monolithic Preference Optimization Without Reference Model.”
arXiv Preprint arXiv:2403.07691.
Horthy, Dex. 2025. 12-Factor Agents: Patterns of Reliable LLM
Applications. Https://github.com/humanlayer/12-factor-agents.
Horvitz, Eric. 1999. “Principles of Mixed-Initiative User
Interfaces.” Proceedings of the SIGCHI Conference on Human
Factors in Computing Systems (CHI), 159–66.
Hsieh, Cheng-Ping, Simeng Sun, Samuel Kriman, et al. 2024. “RULER:
What’s the Real Context Size of Your Long-Context Language
Models?” arXiv Preprint arXiv:2404.06654.
Hu, Edward J., Yelong Shen, Phillip Wallis, et al. 2021.
“LoRA: Low-Rank Adaptation of Large Language
Models.” arXiv Preprint arXiv:2106.09685.
Hu, Shengding, Yuge Tu, Xu Han, et al. 2024.
“MiniCPM: Unveiling the Potential of Small Language Models with
Scalable Training Strategies.” arXiv Preprint
arXiv:2404.06395.
Huang, Jie, Xinyun Chen, Swaroop Mishra, et al. 2024. “Large
Language Models Cannot Self-Correct Reasoning Yet.”
International Conference on Learning Representations.
Huang, Yanping, Youlong Cheng, Ankur Bapna, et al. 2019. “GPipe:
Efficient Training of Giant Neural Networks Using Pipeline
Parallelism.” Advances in Neural Information Processing
Systems 32.
Ilharco, Gabriel, Marco Tulio Ribeiro, Mitchell Wortsman, et al. 2022.
“Editing Models with Task Arithmetic.” arXiv Preprint
arXiv:2212.04089.
ISO/IEC. 2023. ISO/IEC 42001:2023 — Information Technology —
Artificial Intelligence — Management System. International
Organization for Standardization.
Jacobs, Sam Ade, Masahiro Tanaka, Chengming Zhang, et al. 2023.
“DeepSpeed Ulysses: System Optimizations for Enabling Training of
Extreme Long Sequence Transformer Models.” arXiv Preprint
arXiv:2309.14509.
Jeong, Soyeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang, and Jong C.
Park. 2024. “Adaptive-RAG: Learning to Adapt Retrieval-Augmented
Large Language Models Through Question Complexity.”
Proceedings of NAACL-HLT.
Ji, Yichao. 2025. Context Engineering for AI Agents:
Lessons from Building Manus. Manus engineering blog. https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus.
Jiang, Dongfu et al. 2025.
“VerlTool: Towards Holistic Agentic Reinforcement
Learning with Tool Use.” arXiv Preprint
arXiv:2509.01055.
Jiang, Ziheng, Haibin Lin, Yinmin Zhong, et
al. 2024. “MegaScale: Scaling Large Language Model Training
to More Than 10,000 GPUs.” 21st USENIX Symposium on Networked
Systems Design and Implementation (NSDI).
Jimenez, Carlos E., John Yang, et al. 2024.
“SWE-bench: Can Language Models
Resolve Real-World GitHub Issues?” International
Conference on Learning Representations.
Jin, Bowen, Hansi Zeng, Zhenrui Yue, Dong Wang, Hamed Zamani, and Jiawei
Han. 2025. “Search-R1: Training LLMs to Reason and Leverage Search
Engines with Reinforcement Learning.” arXiv Preprint
arXiv:2503.09516.
Jordan, Keller, Yuchen Jin, Vlado Boza, et al. 2024. Muon: An
Optimizer for Hidden Layers in Neural Networks. Https://kellerjordan.github.io/posts/muon/.
Kaplan, Jared et al. 2020. “Scaling
Laws for Neural Language Models.” arXiv Preprint
arXiv:2001.08361.
Kapoor, Sayash, Benedikt Stroebl, Zachary S. Siegel, Nitya Nadgir, and
Arvind Narayanan. 2024. “AI Agents That Matter.” arXiv
Preprint arXiv:2407.01502.
Karpukhin, Vladimir, Barlas Oğuz, Sewon Min, et al. 2020. Dense
Passage Retrieval for Open-Domain Question Answering. https://arxiv.org/abs/2004.04906.
Kazemnejad, Amirhossein, Inkit Padhi, Karthikeyan Natesan Ramamurthy,
Payel Das, and Siva Reddy. 2023. “The Impact of Positional
Encoding on Length Generalization in Transformers.” Advances
in Neural Information Processing Systems.
Kerbl, Bernhard, Georgios Kopanas, Thomas Leimkühler, and George
Drettakis. 2023. “3D Gaussian Splatting for Real-Time
Radiance Field Rendering.” ACM Transactions on Graphics
42 (4).
Khattab, Omar, Arnav Singhvi, Paridhi Maheshwari, et al. 2024.
“DSPy: Compiling Declarative Language Model Calls
into Self-Improving Pipelines.” International Conference on
Learning Representations (ICLR).
Khattab, Omar, and Matei Zaharia. 2020. ColBERT: Efficient and
Effective Passage Search via Contextualized Late Interaction over
BERT. https://arxiv.org/abs/2004.12832.
Kim, Moo Jin, Karl Pertsch, Siddharth Karamcheti,
et al. 2024. “OpenVLA: An Open-Source
Vision-Language-Action Model.” arXiv Preprint
arXiv:2406.09246.
Kingma, Diederik P., and Max Welling. 2014. “Auto-Encoding
Variational Bayes.” International Conference on
Learning Representations (ICLR). https://arxiv.org/abs/1312.6114.
Kleppmann, Martin. 2017. Designing Data-Intensive Applications.
O’Reilly Media.
Kohavi, Ron, Roger Longbotham, Dan Sommerfield, and Randal M. Henne.
2009. “Controlled Experiments on the Web: Survey and Practical
Guide.” Data Mining and Knowledge Discovery 18 (1):
140–81.
Korthikanti, Vijay, Jared Casper, Sangkug Lym, et al. 2022.
“Reducing Activation Recomputation in Large Transformer
Models.” arXiv Preprint arXiv:2205.05198.
Kusupati, Aditya, Gantavya Bhatt, Aniket Rege, et al. 2022.
Matryoshka Representation Learning. https://arxiv.org/abs/2205.13147.
Kwon, Woosuk et al. 2023. “Efficient
Memory Management for Large Language Model Serving with
PagedAttention.” Proceedings of the ACM SIGOPS Symposium on
Operating Systems Principles.
Lee, Harrison, Samrat Phatale, Hassan Mansoor, et al. 2023.
“RLAIF: Scaling Reinforcement Learning from Human
Feedback with AI Feedback.” arXiv Preprint
arXiv:2309.00267.
Lee, John D., and Katrina A. See. 2004. “Trust in Automation:
Designing for Appropriate Reliance.” Human Factors 46
(1): 50–80.
Lee, Katherine, Daphne Ippolito, Andrew Nystrom, et al. 2022.
“Deduplicating Training Data Makes Language Models Better.”
Proceedings of the 60th Annual Meeting of the Association for
Computational Linguistics.
Lepikhin, Dmitry, HyoukJoong Lee, Yuanzhong Xu, et al. 2020.
“GShard: Scaling Giant Models with Conditional
Computation and Automatic Sharding.” arXiv Preprint
arXiv:2006.16668.
Leviathan, Yaniv, Matan Kalman, and Yossi Matias. 2023. “Fast
Inference from Transformers via Speculative Decoding.”
Proceedings of the 40th International Conference on Machine
Learning.
Levine, Sergey, Aviral Kumar, George Tucker, and Justin Fu. 2020.
“Offline Reinforcement Learning: Tutorial, Review, and
Perspectives on Open Problems.” arXiv Preprint
arXiv:2005.01643.
Lewis, Patrick et al. 2020.
“Retrieval-Augmented Generation for Knowledge-Intensive NLP
Tasks.” Advances in Neural Information Processing
Systems.
Li, Kaixin, Ziyang Meng, Hongzhan Lin, et al. 2025.
“ScreenSpot-Pro: GUI Grounding for Professional High-Resolution
Computer Use.” arXiv Preprint arXiv:2504.07981.
Li, Nathaniel et al. 2024. “The WMDP
Benchmark: Measuring and Reducing Malicious Use with Unlearning.”
arXiv Preprint arXiv:2403.03218.
Li, Xuanlin, Kyle Hsu, Jiayuan Gu, et al.
2024. “Evaluating Real-World Robot Manipulation Policies in
Simulation.” arXiv Preprint arXiv:2405.05941.
Li, Yifan, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong
Wen. 2023. “Evaluating Object Hallucination in Large
Vision-Language Models.” Conference on Empirical Methods in
Natural Language Processing (EMNLP).
Li, Yuetai, Xiang Yue, Zhangchen Xu, et al. 2025. “Small Models
Struggle to Learn from Strong Reasoners.” Findings of the
Association for Computational Linguistics: ACL 2025.
Liang, Percy, Rishi Bommasani, Tony Lee, Dimitris
Tsipras, Dilara Soylu, et al. 2022. “Holistic Evaluation of
Language Models.” arXiv Preprint arXiv:2211.09110.
Lightman, Hunter, Vineet Kosaraju, Yura Burda, et al. 2024. “Let’s
Verify Step by Step.” International Conference on Learning
Representations.
Lin, Fangru, Shaoguang Mao, Emanuele La Malfa,
Valentin Hofmann, Adrian de Wynter, et al. 2024. “One
Language, Many Gaps: Evaluating Dialect Fairness and Robustness of Large
Language Models in Reasoning Tasks.” arXiv Preprint
arXiv:2410.11005.
Lin, Ji, Jiaming Tang, Haotian Tang, et al. 2024.
“AWQ: Activation-Aware Weight Quantization for
on-Device LLM Compression and Acceleration.” Proceedings of
Machine Learning and Systems.
Lipman, Yaron, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and
Matt Le. 2023. “Flow Matching for Generative Modeling.”
International Conference on Learning Representations (ICLR). https://arxiv.org/abs/2210.02747.
Little, John D. C. 1961. “A Proof for the Queuing Formula:
L = λW.”
Operations Research 9 (3): 383–87.
Liu, Bo, Yifeng Zhu, Chongkai Gao, et al. 2023.
“LIBERO: Benchmarking Knowledge Transfer for Lifelong
Robot Learning.” Advances in Neural Information Processing
Systems (NeurIPS).
Liu, Haotian, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023.
“Visual Instruction Tuning.” Advances in Neural
Information Processing Systems.
Liu, Hao, Matei Zaharia, and Pieter Abbeel. 2023. “Ring Attention
with Blockwise Transformers for Near-Infinite Context.” arXiv
Preprint arXiv:2310.01889.
Liu, Mingjie, Shizhe Diao, Ximing Lu, et al.
2025. “ProRL: Prolonged Reinforcement Learning
Expands Reasoning Boundaries in Large Language Models.” arXiv
Preprint arXiv:2505.24864.
Liu, Nelson F. et al. 2024. “Lost in
the Middle: How Language Models Use Long Contexts.”
Transactions of the Association for Computational Linguistics.
Liu, Qian, Xiaosen Zheng, Niklas Muennighoff, et
al. 2024. “RegMix: Data Mixture as Regression for Language
Model Pre-Training.” arXiv Preprint arXiv:2407.01492.
Liu, Zichen, Changyu Chen, Wenjun Li, et al. 2025. “Understanding
R1-Zero-Like Training: A Critical Perspective.”
arXiv Preprint arXiv:2503.20783.
Longpre, Shayne, Robert Mahari, Ariel Lee, et
al. 2024. “Consent in Crisis: The Rapid Decline of the AI
Data Commons.” arXiv Preprint arXiv:2407.14933.
Loshchilov, Ilya, and Frank Hutter. 2019. “Decoupled Weight Decay
Regularization.” International Conference on Learning
Representations.
Lu, Enzhe, Zhejun Jiang, Jingyuan Liu, et
al. 2025. “MoBA: Mixture of Block Attention
for Long-Context LLMs.” arXiv Preprint
arXiv:2502.13189.
Lu, Yadong, Jianwei Yang, Yelong Shen, and Ahmed Awadalla. 2025.
“OmniParser for Pure Vision Based GUI Agent.” IEEE/CVF
Conference on Computer Vision and Pattern Recognition (CVPR).
Luo, Xufang et al. 2025. “Agent
Lightning: Train ANY AI Agents with Reinforcement
Learning.” arXiv Preprint arXiv:2508.03680.
Maini, Pratyush, Skyler Seto, He Bai, et al.
2024. “Rephrasing the Web: A Recipe for Compute and Data-Efficient
Language Modeling.” arXiv Preprint arXiv:2401.16380.
Malkov, Yu A., and D. A. Yashunin. 2018. Efficient and Robust
Approximate Nearest Neighbor Search Using Hierarchical Navigable Small
World Graphs. https://arxiv.org/abs/1603.09320.
Mathew, Minesh, Dimosthenis Karatzas, and C. V. Jawahar. 2021.
“DocVQA: A Dataset for VQA on Document Images.”
IEEE/CVF Winter Conference on Applications of Computer Vision
(WACV).
McMahan, Brendan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise
Agüera y Arcas. 2017. “Communication-Efficient Learning of Deep
Networks from Decentralized Data.” Proceedings of the 20th
International Conference on Artificial Intelligence and Statistics
(AISTATS).
Meinke, Alexander et al. 2024.
“Frontier Models Are Capable of in-Context Scheming.”
arXiv Preprint arXiv:2412.04984.
Meng, Yu, Mengzhou Xia, and Danqi Chen. 2024. “SimPO:
Simple Preference Optimization with a Reference-Free Reward.”
Advances in Neural Information Processing Systems 37.
Meta AI. 2025. Agents Rule of Two: A Practical Approach to Secure
Agent Design. Https://ai.meta.com/blog/practical-ai-agent-security/.
METR. 2024. Guidelines for Capability Elicitation. METR
Evaluations.
Mildenhall, Ben, Pratul P. Srinivasan, Matthew Tancik, Jonathan T.
Barron, Ravi Ramamoorthi, and Ren Ng. 2020. “NeRF:
Representing Scenes as Neural Radiance Fields for View
Synthesis.” European Conference on Computer Vision
(ECCV).
Min, Sewon, Xinxi Lyu, Ari Holtzman, et al. 2022. “Rethinking the
Role of Demonstrations: What Makes in-Context Learning Work?”
Proceedings of the 2022 Conference on Empirical Methods in Natural
Language Processing (EMNLP).
Modarressi, Ali, Hanieh Deilamsalehy, Franck Dernoncourt, et al. 2025.
“NoLiMa: Long-Context Evaluation Beyond Literal Matching.”
arXiv Preprint arXiv:2502.05167.
Muennighoff, Niklas, Alexander M. Rush, Boaz Barak,
et al. 2023. “Scaling Data-Constrained Language
Models.” Advances in Neural Information Processing
Systems.
Muennighoff, Niklas, Nouamane Tazi, Loïc Magne, and Nils Reimers. 2023.
MTEB: Massive Text Embedding Benchmark. https://arxiv.org/abs/2210.07316.
Muennighoff, Niklas, Zitong Yang, Weijia Shi, et al. 2025. “S1:
Simple Test-Time Scaling.” arXiv Preprint
arXiv:2501.19393.
Muralidharan, Saurav, Sharath Turuvekere Sreenivas, Raviraj Joshi, et
al. 2024. “Compact Language Models via Pruning and Knowledge
Distillation.” arXiv Preprint arXiv:2407.14679.
Narayanan, Deepak, Mohammad Shoeybi, Jared Casper, et al. 2021.
“Efficient Large-Scale Language Model Training on GPU Clusters
Using Megatron-LM.” Proceedings of the International
Conference for High Performance Computing, Networking, Storage and
Analysis (SC).
National Institute of Standards and Technology. 2023. Artificial
Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1.
NIST.
nostalgebraist. 2020. Interpreting GPT:
The Logit Lens. LessWrong.
Open X-Embodiment Collaboration. 2024. “Open
X-Embodiment: Robotic Learning Datasets and
RT-X Models.” IEEE International Conference on
Robotics and Automation (ICRA).
OpenAI. 2025. “GDPval: Evaluating AI
Model Performance on Economically Valuable Tasks.” arXiv
Preprint arXiv:2510.04374.
Opsahl-Ong, Krista, Michael J. Ryan, Josh Purtell, et al. 2024.
“Optimizing Instructions and Demonstrations for Multi-Stage
Language Model Programs.” Proceedings of the 2024 Conference
on Empirical Methods in Natural Language Processing (EMNLP).
Ouyang, Linke, Yuan Qu, Hongbin Zhou, et al. 2025. “OmniDocBench:
Benchmarking Diverse PDF Document Parsing with Comprehensive
Annotations.” IEEE/CVF Conference on Computer Vision and
Pattern Recognition (CVPR).
Ouyang, Long et al. 2022. “Training
Language Models to Follow Instructions with Human Feedback.”
Advances in Neural Information Processing Systems.
Packer, Charles et al. 2023.
“MemGPT: Towards LLMs as Operating
Systems.” arXiv Preprint arXiv:2310.08560.
Pan, Jiayi et al. 2024. “Training
Software Engineering Agents and Verifiers with
SWE-Gym.” arXiv Preprint arXiv:2412.21139.
Patel, Pratyush, Esha Choukse, Chaojie Zhang, et al. 2024.
“Splitwise: Efficient Generative LLM Inference Using Phase
Splitting.” 2024 ACM/IEEE 51st Annual International Symposium
on Computer Architecture (ISCA), 118–32.
Patterson, David et al. 2025.
“Measuring the Environmental Impact of Delivering AI at Google
Scale.” arXiv Preprint arXiv:2508.15734.
Penedo, Guilherme, Hynek Kydlı́ček, Anton Lozhkov,
et al. 2024. “The FineWeb Datasets: Decanting the Web for
the Finest Text Data at Scale.” arXiv Preprint
arXiv:2406.17557.
Peng, Bowen, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. 2023.
“YaRN: Efficient Context Window Extension of Large Language
Models.” arXiv Preprint arXiv:2309.00071.
Pertsch, Karl, Kyle Stachowicz, Brian Ichter, et
al. 2025. “FAST: Efficient Action Tokenization
for Vision-Language-Action Models.” arXiv Preprint
arXiv:2501.09747.
Press, Ofir, and Lior Wolf. 2017. “Using the Output Embedding to
Improve Language Models.” Proceedings of the Conference of
the European Chapter of the Association for Computational
Linguistics.
Qi, Charles R., Hao Su, Kaichun Mo, and Leonidas J. Guibas. 2017.
“PointNet: Deep Learning on Point Sets for
3D Classification and Segmentation.” IEEE
Conference on Computer Vision and Pattern Recognition (CVPR).
Radford, Alec, Jong Wook Kim, Chris Hallacy, et al. 2021.
“Learning Transferable Visual Models from Natural Language
Supervision.” International Conference on Machine
Learning.
Radford, Alec, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and
Ilya Sutskever. 2019. “Language Models Are Unsupervised Multitask
Learners.” OpenAI Technical Report.
Rafailov, Rafael et al. 2023. “Direct
Preference Optimization: Your Language Model Is Secretly a Reward
Model.” Advances in Neural Information Processing
Systems.
Raffel, Colin, Noam Shazeer, Adam Roberts, et
al. 2020. “Exploring the Limits of Transfer Learning with a
Unified Text-to-Text Transformer.” Journal of Machine
Learning Research 21.
Rajbhandari, Samyam, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020.
“ZeRO: Memory Optimizations Toward Training Trillion Parameter
Models.” arXiv Preprint arXiv:1910.02054.
Rasmussen, Preston et al. 2025.
“Zep: A Temporal Knowledge Graph Architecture for
Agent Memory.” arXiv Preprint arXiv:2501.13956.
Richardson, Chris. 2019. Transactional Outbox. Https://microservices.io/patterns/data/transactional-outbox.html.
Robertson, Stephen, and Hugo Zaragoza. 2009. “The Probabilistic
Relevance Framework: BM25 and Beyond.” Foundations and Trends
in Information Retrieval 3 (4): 333–89.
Ross, Stéphane, Geoffrey J. Gordon, and J. Andrew Bagnell. 2011.
“A Reduction of Imitation Learning and Structured Prediction to
No-Regret Online Learning.” International Conference on
Artificial Intelligence and Statistics.
Sardana, Nikhil, Jacob Portes, Sasha Doubov, and Jonathan Frankle. 2024.
“Beyond Chinchilla-Optimal: Accounting for Inference in Language
Model Scaling Laws.” International Conference on Machine
Learning.
Schaeffer, Rylan, Brando Miranda, and Sanmi Koyejo. 2023. “Are
Emergent Abilities of Large Language Models a Mirage?”
Advances in Neural Information Processing Systems.
Schulman, John, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg
Klimov. 2017. “Proximal Policy Optimization Algorithms.”
arXiv Preprint arXiv:1707.06347.
Sennrich, Rico, Barry Haddow, and Alexandra Birch. 2016. “Neural
Machine Translation of Rare Words with Subword Units.”
Proceedings of the Annual Meeting of the Association for
Computational Linguistics.
Shao, Zhihong, Peiyi Wang, Qihao Zhu, et al.
2024. “DeepSeekMath: Pushing the Limits of
Mathematical Reasoning in Open Language Models.” arXiv
Preprint arXiv:2402.03300.
Shazeer, Noam. 2019. “Fast Transformer Decoding: One Write-Head Is
All You Need.” arXiv Preprint arXiv:1911.02150.
Shazeer, Noam. 2020. “GLU Variants Improve Transformer.”
arXiv Preprint arXiv:2002.05202.
Shazeer, Noam, Azalia Mirhoseini, Krzysztof Maziarz, et al. 2017.
“Outrageously Large Neural Networks: The Sparsely-Gated
Mixture-of-Experts Layer.” arXiv Preprint
arXiv:1701.06538.
Sheng, Ying, Shiyi Cao, Dacheng Li, et al. 2024.
“S-LoRA: Serving Thousands of Concurrent
LoRA Adapters.” Proceedings of Machine Learning
and Systems.
Shinn, Noah et al. 2023. “Reflexion:
Language Agents with Verbal Reinforcement Learning.” Advances
in Neural Information Processing Systems.
Shoeybi, Mohammad, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared
Casper, and Bryan Catanzaro. 2019. “Megatron-LM: Training
Multi-Billion Parameter Language Models Using Model Parallelism.”
arXiv Preprint arXiv:1909.08053.
Shokri, Reza, Marco Stronati, Congzheng Song, and Vitaly Shmatikov.
2017. “Membership Inference Attacks Against Machine Learning
Models.” IEEE Symposium on Security and Privacy.
Shumailov, Ilia, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross
Anderson, and Yarin Gal. 2024. “AI Models Collapse When Trained on
Recursively Generated Data.” Nature 631: 755–59.
Simmons, Joseph P., Leif D. Nelson, and Uri Simonsohn. 2011.
“False-Positive Psychology: Undisclosed Flexibility in Data
Collection and Analysis Allows Presenting Anything as
Significant.” Psychological Science 22 (11).
Skelton, Matthew, and Manuel Pais. 2019. Team Topologies: Organizing
Business and Technology Teams for Fast Flow. IT Revolution Press.
Smith, Reid G. 1980. “The Contract Net Protocol: High-Level
Communication and Control in a Distributed Problem Solver.”
IEEE Transactions on Computers C-29 (12): 1104–13.
Snell, Charlie, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024.
“Scaling LLM Test-Time Compute Optimally Can Be More
Effective Than Scaling Model Parameters.” arXiv Preprint
arXiv:2408.03314.
Sorensen, Tyler, and Heidy Khlaaf. 2024.
LeftoverLocals: Listening to LLM Responses
Through Leaked GPU Local Memory. Trail of Bits
Research.
Su, Jianlin, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng
Liu. 2021. “RoFormer: Enhanced Transformer with Rotary Position
Embedding.” arXiv Preprint arXiv:2104.09864.
Sumers, Theodore R., Shunyu Yao, Karthik Narasimhan, and Thomas L.
Griffiths. 2023. “Cognitive Architectures for Language
Agents.” arXiv Preprint arXiv:2309.02427.
Szot, Andrew, Alexander Clegg, Eric Undersander, et
al. 2021. “Habitat 2.0: Training Home Assistants to
Rearrange Their Habitat.” Advances in Neural Information
Processing Systems (NeurIPS).
Tam, Zhi Rui, Cheng-Kuang Wu, Yi-Lin Tsai, Chieh-Yen Lin, Hung-yi Lee,
and Yun-Nung Chen. 2024. Let Me Speak Freely? A Study on the Impact
of Format Restrictions on Performance of Large Language Models. https://arxiv.org/abs/2408.02442.
Templeton, Adly et al. 2024. “Scaling
Monosemanticity: Extracting Interpretable Features from Claude 3
Sonnet.” Transformer Circuits Thread.
Tobin, Josh, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba,
and Pieter Abbeel. 2017. “Domain Randomization for Transferring
Deep Neural Networks from Simulation to the Real World.”
IEEE/RSJ International Conference on Intelligent Robots and Systems
(IROS).
Turpin, Miles et al. 2023. “Language
Models Don’t Always Say What They Think: Unfaithful Explanations in
Chain-of-Thought Prompting.” Advances in Neural Information
Processing Systems.
.txt team. 2024. Say What You Mean: A
Response to “Let Me Speak Freely”. Https://blog.dottxt.co/say-what-you-mean.html.
Vassilev, Apostol, Alina Oprea, Alie Fordyce, Hyrum
Anderson, et al. 2025. Adversarial Machine Learning: A
Taxonomy and Terminology of Attacks and Mitigations (NIST AI
100-2e2025). National Institute of Standards; Technology. https://doi.org/10.6028/NIST.AI.100-2e2025.
Vaswani, Ashish et al. 2017.
“Attention Is All You Need.” Advances in Neural
Information Processing Systems.
Wang, Guanzhi et al. 2023. “Voyager:
An Open-Ended Embodied Agent with Large Language Models.”
arXiv Preprint arXiv:2305.16291.
Wang, Peiyi, Lei Li, Zhihong Shao, et al. 2024. “Math-Shepherd:
Verify and Reinforce LLMs Step-by-Step Without Human
Annotations.” Annual Meeting of the Association for
Computational Linguistics.
Wang, Xiaohan, Yuhui Zhang, Orr Zohar, and Serena Yeung-Levy. 2024.
“VideoAgent: Long-Form Video Understanding with Large
Language Model as Agent.” European Conference on Computer
Vision (ECCV). https://arxiv.org/abs/2403.10517.
Wang, Xuezhi, Jason Wei, Dale Schuurmans, et al. 2023.
“Self-Consistency Improves Chain of Thought Reasoning in Language
Models.” International Conference on Learning
Representations.
Wang, Yizhong, Yeganeh Kordi, Swaroop Mishra, et al. 2023.
“Self-Instruct: Aligning Language Models with
Self-Generated Instructions.” Proceedings of the 61st Annual
Meeting of the Association for Computational Linguistics.
Wang, Zihan et al. 2025.
“RAGEN: Understanding Self-Evolution in
LLM Agents via Multi-Turn Reinforcement Learning.”
arXiv Preprint arXiv:2504.20073.
Wei, Alexander, Nika Haghtalab, and Jacob Steinhardt. 2023.
“Jailbroken: How Does LLM Safety Training Fail?”
Advances in Neural Information Processing Systems.
Wei, Jason, Yi Tay, Rishi Bommasani, Colin Raffel,
et al. 2022. “Emergent Abilities of Large Language
Models.” Transactions on Machine Learning Research.
Wei, Jason, Xuezhi Wang, Dale Schuurmans, et al. 2022.
“Chain-of-Thought Prompting Elicits Reasoning in Large Language
Models.” Advances in Neural Information Processing
Systems 35.
Weij, Teun van der et al. 2024. “AI
Sandbagging: Language Models Can Strategically Underperform on
Evaluations.” arXiv Preprint arXiv:2406.07358.
Willard, Brandon T., and Rémi Louf. 2023. Efficient Guided
Generation for Large Language Models. https://arxiv.org/abs/2307.09702.
Williams, Grady, Nolan Wagener, Brian Goldfain, et al. 2017.
“Information Theoretic MPC for Model-Based
Reinforcement Learning.” IEEE International Conference on
Robotics and Automation (ICRA).
Williams, Ronald J. 1992. “Simple Statistical Gradient-Following
Algorithms for Connectionist Reinforcement Learning.” Machine
Learning 8: 229–56.
Williams, Samuel, Andrew Waterman, and David Patterson. 2009.
“Roofline: An Insightful Visual Performance Model for Multicore
Architectures.” Communications of the ACM.
Willison, Simon. 2023. The Dual LLM Pattern for Building AI
Assistants That Can Resist Prompt Injection. Https://simonwillison.net/2023/Apr/25/dual-llm-pattern/.
Wu, Di et al. 2024.
“LongMemEval: Benchmarking Chat Assistants on
Long-Term Interactive Memory.” arXiv Preprint
arXiv:2410.10813.
Xiao, Guangxuan, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and
Song Han. 2023. “SmoothQuant: Accurate and Efficient Post-Training
Quantization for Large Language Models.” Proceedings of the
40th International Conference on Machine Learning.
Xiao, Guangxuan, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis.
2024. “Efficient Streaming Language Models with Attention
Sinks.” International Conference on Learning
Representations.
Xie, Sang Michael, Hieu Pham, Xuanyi Dong, et
al. 2023. “DoReMi: Optimizing Data Mixtures Speeds up
Language Model Pretraining.” Advances in Neural Information
Processing Systems.
Xie, Tianbao, Danyang Zhang, Jixuan Chen, et al. 2024. OSWorld:
Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer
Environments.
Yadav, Prateek, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit
Bansal. 2023. “TIES-Merging: Resolving Interference
When Merging Models.” Advances in Neural Information
Processing Systems.
Yan, Shi-Qi, Jia-Chen Gu, Yun Zhu, and Zhen-Hua Ling. 2024.
“Corrective Retrieval Augmented Generation.” arXiv
Preprint arXiv:2401.15884.
Yang, Greg, Edward J. Hu, Igor Babuschkin, et al. 2022. “Tensor
Programs v: Tuning Large Neural Networks via Zero-Shot Hyperparameter
Transfer.” Advances in Neural Information Processing
Systems 35.
Yang, Jianwei, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, and Jianfeng
Gao. 2023. “Set-of-Mark Prompting Unleashes Extraordinary Visual
Grounding in GPT-4V.” arXiv Preprint arXiv:2310.11441.
Yang, John, Carlos E. Jimenez, Alexander Wettig, et al. 2024.
“SWE-agent: Agent-Computer Interfaces
Enable Automated Software Engineering.” Advances in Neural
Information Processing Systems.
Yang, Songlin, Jan Kautz, and Ali Hatamizadeh. 2024. “Gated Delta
Networks: Improving Mamba2 with Delta Rule.” arXiv Preprint
arXiv:2412.06464.
Yao, Shunyu et al. 2023. “ReAct:
Synergizing Reasoning and Acting in Language Models.”
International Conference on Learning Representations.
Yao, Shunyu et al. 2024. “τ-bench:
A Benchmark for Tool-Agent-User Interaction in Real-World
Domains.” arXiv Preprint arXiv:2406.12045.
Yao, Shunyu, Dian Yu, Jeffrey Zhao, et al. 2023. “Tree of
Thoughts: Deliberate Problem Solving with Large Language Models.”
Advances in Neural Information Processing Systems 36.
Young, John W. 1974. “A First Order Approximation to the Optimum
Checkpoint Interval.” Communications of the ACM 17 (9):
530–31.
Yu, Gyeong-In, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and
Byung-Gon Chun. 2022. “Orca: A Distributed Serving System for
Transformer-Based Generative Models.” 16th USENIX Symposium
on Operating Systems Design and Implementation (OSDI 22), 521–38.
Yu, Qiying, Zheng Zhang, Ruofei Zhu, Yufeng Yuan,
et al. 2025. “DAPO: An Open-Source
LLM Reinforcement Learning System at Scale.”
arXiv Preprint arXiv:2503.14476.
Yuan, Jingyang, Huazuo Gao, Damai Dai, et
al. 2025. “Native Sparse Attention: Hardware-Aligned and
Natively Trainable Sparse Attention.” arXiv Preprint
arXiv:2502.11089.
Yue, Yang, Zhiqi Chen, Rui Lu, et al. 2025. “Does Reinforcement
Learning Really Incentivize Reasoning Capacity in LLMs
Beyond the Base Model?” arXiv Preprint arXiv:2504.13837.
Zeghidour, Neil, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco
Tagliasacchi. 2021. “SoundStream: An End-to-End
Neural Audio Codec.” IEEE/ACM Transactions on Audio, Speech,
and Language Processing. https://arxiv.org/abs/2107.03312.
Zhai, Xiaohua, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer.
2023. “Sigmoid Loss for Language Image Pre-Training.”
IEEE/CVF International Conference on Computer Vision (ICCV).
Zhang, Andy K. et al. 2024. “Cybench:
A Framework for Evaluating Cybersecurity Capabilities and Risks of
Language Models.” arXiv Preprint arXiv:2408.08926.
Zhang, Biao, and Rico Sennrich. 2019. “Root Mean Square Layer
Normalization.” Advances in Neural Information Processing
Systems.
Zhang, Qizheng et al. 2025. “Agentic
Context Engineering: Evolving Contexts for Self-Improving Language
Models.” arXiv Preprint arXiv:2510.04618.
Zhao, Yanli, Andrew Gu, Rohan Varma, et al. 2023. “PyTorch FSDP:
Experiences on Scaling Fully Sharded Data Parallel.”
Proceedings of the VLDB Endowment 16 (12): 3848–60.
Zheng, Lianmin, Wei-Lin Chiang, Ying Sheng, et
al. 2023. “Judging LLM-as-a-Judge with
MT-Bench and Chatbot Arena.”
Advances in Neural Information Processing Systems.
Zheng, Lianmin, Liangsheng Yin, Zhiqiang Xie, et al. 2024.
“SGLang: Efficient Execution of Structured Language Model
Programs.” Advances in Neural Information Processing
Systems.
Zheng, Mingqian, Jiaxin Pei, Lajanugen Logeswaran, Moontae Lee, and
David Jurgens. 2024. “When ‘a Helpful Assistant’ Is
Not Really Helpful: Personas in System Prompts Do Not Improve
Performances of Large Language Models.” Findings of the
Association for Computational Linguistics: EMNLP 2024.
Zhong, Yinmin, Shengyu Liu, Junda Chen, et al. 2024. “DistServe:
Disaggregating Prefill and Decoding for Goodput-Optimized Large Language
Model Serving.” 18th USENIX Symposium on Operating Systems
Design and Implementation (OSDI 24), 193–210.
Zhou, Chunting, Pengfei Liu, Puxin Xu, et
al. 2023. “LIMA: Less Is More for
Alignment.” Advances in Neural Information Processing
Systems 36.
Zhou, Denny, Nathanael Schärli, Le Hou, Jason Wei,
et al. 2023. “Least-to-Most Prompting Enables Complex
Reasoning in Large Language Models.” International Conference
on Learning Representations.
Zhou, Shuyan et al. 2024.
“WebArena: A Realistic Web Environment for Building
Autonomous Agents.” International Conference on Learning
Representations.
Zoph, Barret, Irwan Bello, Sameer Kumar, et al. 2022.
“ST-MoE: Designing Stable and Transferable Sparse
Expert Models.” arXiv Preprint arXiv:2202.08906.
Zou, Wei, Runpeng Geng, Binghui Wang, and Jinyuan Jia. 2025.
“PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented
Generation of Large Language Models.” USENIX Security
Symposium.