An AI-enabled Digital Research Assistant for the Legislación Mexicana Corpus
DOI:
https://doi.org/10.5860/ital.v45i3.17711Abstract
Large-scale historical corpora present persistent challenges for research, particularly in relation to the navigation, interpretation, and extraction of meaningful information across extensive and heterogeneous collections. The Legislación Mexicana corpus, a 42-volume compilation of legal dispositions spanning more than two centuries, exemplifies these challenges due to its size, structural complexity, and evolving terminology. Traditional approaches to working with such corpora have relied on close reading and the use of indexes. While effective within defined limits, these methods constrain the scope of inquiry and require significant time and expertise to produce meaningful results. This article presents the development of LegMexIA, an AI-enabled research assistant designed to extend interaction with the corpus beyond conventional retrieval methods. The system integrates traditional text retrieval and retrieval-augmented generation via a coordinated architecture that distinguishes between different types of user queries. Structured queries are processed using Elasticsearch and BM25 ranking, while exploratory and interpretive queries are routed through a retrieval-augmented generation workflow, where relevant text fragments are retrieved using vector similarity and assembled into contextual inputs for a language model. A key aspect of the system is the agentic decision layer that analyzes user intent and dynamically selects the appropriate processing strategy, in alignment with established reference practices in academic libraries and in a way that maintains the user’s role in verification. The article also addresses key considerations related to prompt design, technological sustainability, institutional constraints, and the implications of public deployment.
References
ColmexBDCV, “legmex_prompts,” GitHub repository, https://github.com/ColmexBDCV/legmex_prompts.
Franco Moretti, “Conjectures on World Literature,” New Left Review 1 (2000): 54–68, https://newleftreview.org/issues/ii1/articles/franco-moretti-conjectures-on-world-literature.pdf.
Frank Boateng, “The Transformative Potential of Generative AI in Academic Library Access Services: Opportunities and Challenges,” Information Services and Use 45, nos. 1–2 (2025): 140–47, https://journals.sagepub.com/doi/10.1177/18758789251332800.
Hans Peter Luhn, “A Statistical Approach to Mechanized Encoding and Searching of Literary Information,” IBM Journal of Research and Development 1, no. 4 (1957): 309–7, https://ieeexplore.ieee.org/abstract/document/5392697.
Kathleen Kluegel, Catherine Sheldrick Ross, Jana Ronan, Kathleen Kern, and David Tyckoson, “The Reference Interview: Connecting in Person and in Cyberspace: Presentations and Responses from the RUSA President’s Program, 2002 ALA Annual Conference, Atlanta, June 17, 2002,” Reference & User Services Quarterly 43, no. 1 (2003): 37–51, https://research.ebsco.com/c/xwhikx/search/details/hqyzhc63ov.
Keerthana Murugaraj, Salima Lamsiyah, Marten During, and Martin Theobald, “Topic-RAG for Historical Newspapers: Enhancing Information Retrieval in Humanities Research through Topic-Based Retrieval-Augmented Generation,” Computational Humanities Research 1 (2025): e15, https://doi.org/10.1017/chr.2025.10018.
Library Innovation Lab, Harvard Law School, “Open French Law RAG,” https://lil.law.harvard.edu/open-french-law-rag/.
“LRAGE: Legal Retrieval Augmented Generation Evaluation Tool,” arXiv, 2025.
Nicholas Pipitone and Ghita Houir Alami, “LegalBench RAG: A Benchmark for Retrieval Augmented Generation in the Legal Domain,” arXiv, August 19, 2024, https://arxiv.org/abs/2408.10343.
P. Kelly, J. Schild, and A. Jafari, “FolkRAG: a Retrieval-Augmented Generation System for Cultural Heritage Materials,” Neural Computing and Applications 37 (2025): 20281–97, https://doi.org/10.1007/s00521-025-11455-4.
Patrick Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” Advances in Neural Information Processing Systems 33 (2020): 9459–74, https://dl.acm.org/doi/abs/10.5555/3495724.3496517.
Robert L. Collison, Indexing Books: A Manual of Basic Principles (London: Ernest Benn Limited, 1962), 26–27, 34–35, 50.
Stephen E. Robertson and Hugo Zaragoza, “The Probabilistic Relevance Framework: BM25 and Beyond,” Foundations and Trends in Information Retrieval 3, no. 4 (2009): 333–89.
Vannevar Bush, “As We May Think,” The Atlantic Monthly 176, no. 1 (1945): 101–8, https://www.ias.ac.in/article/fulltext/reso/005/11/0094-0103.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Alberto Martinez, Rodrigo Cuellar Hidalgo, Javier Cisneros

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Authors that submit to Information Technology and Libraries agree to the Copyright Notice.