Why data is the foundation of all legal AI

When a legal practitioner queries EuroLegalBot, attention naturally focuses on the chatbot and its ability to provide a relevant response quickly. However, the quality of that response depends first and foremost on the quality of the knowledge on which the system is based.

Designed to assist professionals involved in European arrest warrant procedures, EuroLegalBot aims to provide practical, accessible and reliable support in a particularly complex area of law. By utilising natural language processing, machine learning techniques and a structured knowledge base, the system is designed to facilitate access to legal information, direct users to relevant sources and help them better understand the applicable procedural requirements. Its role is not to replace the analysis or judgement of legal practitioners, but to provide them with a support tool based on identified and organised legal sources.

In the legal field, artificial intelligence does not create the law. It draws on existing sources, which it must be able to identify, organise and present in a coherent manner. Building a reliable dataset is therefore a crucial stage of the project. This is precisely the aim of the Dataset deliverable: to transform a complex collection of legal texts, court decisions and procedural documents into a knowledge base that can be utilised by the system.

This work addresses a challenge well known to practitioners in the field of judicial cooperation: relevant information is often scattered across European texts, national legislation, case law and practical guidance. For a legal assistant to be able to find a useful answer quickly, these various sources must first be brought together, organised and linked. This article presents the key findings of the ‘Dataset’ deliverable, which is appended to this publication so that readers wishing to explore certain technical or methodological aspects in greater depth can consult the reference document directly.

From source to knowledge: the aim of the dataset

The ‘Dataset’ deliverable (see attached doucment) is not merely a compilation of documents. Its aim is to create a structured documentation environment that will enable the future chatbot to make consistent use of legal resources relating to the European arrest warrant.

This process involves several successive steps: identifying the sources, collecting them, classifying them, organising them and preparing them for future use. The data collected comes from European and national institutional sources and is converted into formats that can be processed by artificial intelligence tools. It is also cleaned, validated and enriched to ensure its quality and relevance.

Each document included in the corpus is accompanied by descriptive information that identifies its institutional origin, the relevant level of jurisdiction and its territorial scope. This structure is designed to facilitate the organisation of a corpus comprising several thousand legal resources from various European legal systems.

The challenge, therefore, is not merely to compile documents, but to transform a vast amount of disparate legal information into a structured, usable and evolving knowledge base capable of supporting the future legal assistance features developed as part of the EuroLegalBot project.

Why was it necessary to compile 6,339 legal documents?

One of the most significant outcomes of the work carried out is the creation of an initial corpus comprising 6,339 legal documents.

This collection brings together materials from various institutional sources and covers several levels of legislation. The documents have been categorised according to their origin, the jurisdiction to which they relate and their territorial scope. This systematic organisation makes it possible to structure a particularly extensive body of documentation whilst ensuring consistent navigation within the resources.

Behind this figure lies a considerable amount of work involving selection, collation and organisation. For a practitioner faced with a question relating to a European arrest warrant, the answer is rarely to be found in a single source. It may require consulting several levels of information: the applicable European texts, their transposition into national law, relevant case law and the available procedural resources. The compilation of a corpus of 6,339 documents precisely reflects this documentary complexity, which the project aims to make more accessible.

The information does not come from a single source but from a complex ecosystem of documents that reflects the reality of European judicial cooperation.

The aim is twofold:

  • to ensure sufficiently comprehensive documentary coverage;
  • to enable future legal aid tools to make effective use of the information.

What documentary sources make up the dataset?

To build EuroLegalBot’s knowledge base, the consortium members compiled 6,339 documents from European and national institutional sources. This corpus includes, in particular, resources from the Luxembourg Investigating Chamber, the Polish Supreme Court, the Italian Court of Cassation, the EUR-Lex portal, Spanish public sources and numerous national high courts.

Among the most significant collections of documents are those from the Luxembourg Investigating Chamber (777 documents), the ‘European Justice’ directory containing resources of European scope (713 documents), the Supreme Court of Poland (674 documents), Spanish public resources (611 documents), the EUR-Lex portal dedicated to European Union legislation (528 documents) and the Italian Court of Cassation (487 documents). The corpus also includes numerous decisions from German higher courts as well as other national judicial institutions.

This diversity of sources is one of the Dataset’s key strengths. It makes it possible to bring together, within a single documentary environment, legislative texts, court decisions, case law compendia and institutional resources from several European legal systems. The future system will thus be able to draw on a knowledge base that reflects the diversity of practices and sources utilised in the context of European judicial cooperation.

The value of the dataset lies not only in its volume but also in the diversity of the sources from which it is drawn. The European Arrest Warrant is an instrument at the intersection of several legal systems: it involves EU law, national legislation and the courts’ interpretation of these. The inclusion, within a single corpus, of documents from a wide range of European and national institutions thus reflects this complex legal reality and provides a knowledge base covering a broad spectrum of situations likely to be encountered by legal practitioners.

The role of case law in building the knowledge base

Case law plays a central role in the development of legal knowledge. Unlike statutory texts, court decisions provide interpretative guidance that is essential to the understanding and application of the law. They enable us to understand how the rules are actually applied in real-life situations and to bridge the gap between legal knowledge and the practical realities faced by legal practitioners.

As part of the development of the dataset, we have selected an initial set of 92 court cases intended to form the case-law foundation of the knowledge base. This selection covers nearly twenty years of developments in case law relating to the European Arrest Warrant, from the Advocaten voor de Wereld judgment handed down in 2007 to the most recent decisions included in 2025. The selected cases cover, in particular, issues such as: fundamental rights, double criminality, the ne bis in idem principle, conditions of detention, the independence of judicial authorities, procedural time limits, safeguards for wanted persons, and the conditions for surrender between Member States.

The incorporation of this case law enables the Dataset to go beyond a purely documentary approach. The knowledge base does not merely compile applicable texts; it also incorporates their interpretation and application by the competent courts. This approach helps to establish links between questions that users are likely to ask and situations already examined by European courts, thereby enhancing the relevance and contextualisation of the knowledge utilised within EuroLegalBot.

Data validation, quality and reliability

The value of a legal system supported by artificial intelligence depends directly on the quality of the data it uses.

The work presented in the dataset places particular emphasis on the validation of the corpus. The documents undergo verification, cleaning and organisation processes designed to improve their consistency and usability.

This approach has several objectives:

  • to reduce inconsistencies in documentation;
  • to improve the quality of future research;
  • to ensure the relevance of the results obtained;
  • to enhance the reliability of the knowledge environment.

In a legal context, where the accuracy of information is essential, this step is an indispensable prerequisite for any development based on artificial intelligence.

The dataset as a knowledge infrastructure

Beyond the documents themselves, the dataset should be understood as a genuine knowledge infrastructure.

Its function is not merely to store information, but to organise the relationships between different legal resources. The classification and structuring mechanisms put in place make it possible to create an environment in which texts, decisions and other documents can be retrieved, compared and utilised in a coherent manner.

In order to ensure that the corpus remains up to date, we have set up an automated data collection mechanism that runs weekly, drawing on selected institutional sources. This regular updating enables new decisions and newly published legal instruments to be incorporated gradually, whilst maintaining the consistency of the database over time. The dataset is therefore not intended to be a static snapshot of the law applicable to the European arrest warrant, but rather a knowledge base capable of evolving in step with publications from the relevant courts and institutions.

In practical terms, when a user searches for information relating to a European Arrest Warrant procedure, the aim is not merely to access a single document, but a coherent body of knowledge combining the various applicable texts, the available documentary resources and the relevant court decisions. The true value of the dataset lies in this ability to organise the relationships between sources.

This infrastructure forms the foundation upon which the project’s future developments are built. It represents EuroLegalBot’s information assets and is one of its most important assets.

Thanks to this structuring work, legal knowledge becomes usable on a large scale and can be made available to practitioners via conversational interfaces.

Conclusion: Structured knowledge – a prerequisite for AI-powered legal assistance

The dataset illustrates a reality that is often overlooked in artificial intelligence projects: the quality of a legal assistant depends, above all, on the quality of the knowledge it draws upon.

Before the chatbot could answer a question, it was necessary to identify, collect, organise and classify several thousand documents from various levels of the European legal landscape. This structuring work is a prerequisite for any reliable retrieval of information.

When a legal practitioner queries EuroLegalBot tomorrow, they will not be consulting directly the 6,339 documents that currently make up the project’s corpus. Nevertheless, each of these resources will contribute to the quality of the response received. The dataset thus emerges as the unobtrusive yet essential link connecting European legal sources to their practical application in the context of judicial cooperation.