FASTMIND
FastMind: Fast and Automatic Scientific Text Mining for Integration and Normalisation into Databases
FastMind is an AI-driven data-mining pipeline designed to extract complex scientific information automatically from very large collections of structured documents.
Scientific knowledge is expanding much faster than researchers’ capacity to read, extract and synthesise it. Thousands of relevant studies may exist for a single ecological question, each containing valuable information embedded in text, tables and supplementary material.
FastMind was developed to overcome this bottleneck by combining Large Language Models (LLMs), automated document processing and ecological expertise to transform structured scientific literature into standardized, accurate machine-readable data.
Its first large-scale application is the study of the ecological impacts of biological invasions, through the InvaPact project.
The challenge: scientific information at a scale humans can no longer process
Modern ecology faces a paradox: we have more scientific information than ever before, but much of it remains extremely difficult to synthesise.
Traditional systematic reviews and meta-analyses require researchers to read publications individually and manually extract the relevant information. This remains indispensable for many scientific questions, but it becomes prohibitively expensive when tens of thousands of studies and dozens of variables are involved.

For FastMind’s first application, approximately 55 complex ecological parameters must be extracted from nearly 14,000 scientific publications.
At an estimated four hours of double-blind manual extraction per publication, completing the same task using a conventional systematic-review protocol would require more than 113,000 hours of scientific work — approximately 70 full-time-equivalent years. And the quality of the extraction would invariably go decreasing, with rates of errors unavoidably exploding when experts become tired with the extractions.
FastMind is therefore not intended to replace scientific expertise. It enables expert-designed data extraction to operate at a scale that would otherwise be practically inaccessible.
How FastMind works: LLMs guided by ecological expertise
FastMind combines modern Large Language Models with a scientific workflow explicitly designed around predefined ecological concepts and variables.
Rather than asking an AI system to produce an unconstrained summary of a publication, FastMind asks precise scientific questions and transforms the answers into a standardized database structure.
The general workflow includes:
1 – Identification and preparation of scientific documents
Relevant publications are collected, converted into formats that can be processed automatically, and organised for systematic analysis.
2 – AI-assisted extraction
Large Language Models analyse the documents and extract predefined ecological variables according to detailed scientific instructions.
3 – Standardisation and structuring
Extracted information is translated into homogeneous descriptors and database fields so that information originating from thousands of different studies can be compared.
4 – Quality control and validation
Automated extraction is benchmarked against data independently extracted and validated by researchers. Steps 2-3-4 are repeated until a satisfactory accuracy rate is reached.
5 – Scientific analysis
Once structured, the resulting database can be explored using conventional statistical, macroecological and meta-analytical approaches.
This architecture makes the system adaptable: the extraction framework can be modified to address different ecological questions without rebuilding the entire pipeline from scratch.
A large-scale test: quantifying biological invasions
FastMind’s first major application is InvaPact, an international research programme developing a standardized quantitative metric for the ecological impacts of invasive alien species.
Biological invasion studies provide a particularly demanding test for automated scientific extraction. Relevant information may concern species, ecosystems, mechanisms of impact, spatial and temporal scale, ecological organisation, severity and many other variables that are rarely reported using identical terminology.
FastMind was designed to extract approximately 55 parameters for each documented invasion impact from a corpus of nearly 14,000 scientific publication.
This effort has generated approximately 160,000 structured ecological impact records, forming the core of the InvaPact database.
The database is designed to support global analyses of invasion impacts across species, ecosystems, geographic regions and ecological mechanisms, and to provide the empirical foundation for the new InvaPact quantitative impact metric.

Validation: measuring AI performance against expert-curated data
For scientific applications, speed alone is not sufficient: extracted information must also be reliable.
FastMind has therefore been evaluated against Golden Standard publications, for which the same ecological information was independently extracted and validated manually by researchers.
Across the complex variables evaluated, the FastMind pipeline achieved an overall accuracy much higher than human experts. This combination of automated benchmarking and expert validation is designed to identify systematic errors, improve extraction instructions and progressively strengthen the reliability of the extractions.
Beyond InvaPact: towards a general scientific data-mining framework
Although biological invasions provide FastMind’s first large-scale application, the underlying challenge is universal across science: enormous quantities of potentially valuable information remain trapped in unstructured documents.
The long-term objective is therefore to make the FastMind methodology transferable to other ecological and scientific questions.
Current developments include:
– extraction from scientific literature written in languages other than English, including Chinese, French or Spanish;
– improved processing of complex document structures;
– application to grey literature and unpublished reports;
– adaptation to other ecological research questions requiring large standardized databases.
The extension to unpublished environmental documentation is giving rise to two related initiatives:
– InvaGrey, focused on extracting invasion information from grey literature, such as report produced by NGOs and environmental agencies;
– InvaCorp, which explores how the FastMind–InvaPact approach can be adapted to companies’ environmental impact assessments, monitoring reports and other internal environmental documents.
FastMind therefore serves both as a scientific tool for the InvaPact programme and as a broader methodological platform for transforming large bodies of unstructured ecological knowledge into quantitative data.

A combination of artificial and human intelligence
FastMind is developed by researchers at the CNRS, Université Paris-Saclay and MNHN through a collaboration between specialists in ecology, biological invasions, artificial intelligence and scientific data management. The project builds on the experience acquired through large international synthesis projects such as InvaCost and InvaPact, and on an extensive international network of over 160 invasion scientists from 55 countries.
Its central principle is that artificial intelligence and scientific expertise should not be opposed: AI provides the scale, while scientists define the questions, the concepts, the validation procedures and the interpretation. Our LLM are fully supervised in the sense that each step is validated by experts, most of the time several times and independently. This combination makes it possible to address scientific questions that would otherwise require decades of manual data extraction.
FastMind has benefited from a grant by the OFB, ending soon. Do not hesitate to contact us if you want to support the continuation of this project.
