Contractzlab Research

Making legal AI a scientific discipline, not a marketing promise.

Contractzlab Research is the AI laboratory that designs, trains, and evaluates the models driving the Contractzlab platform. Our mission: produce rigorous research in law, compliance, and risk management, and turn it into reliable systems for demanding organizations.

View our publications
3
active research focus areas
436
legal questions in our latest benchmark
2
continents covered by evaluation data
100 %
in-house research
Our Lab

A research team backed by a production-grade product

Contractzlab Research stems from a simple observation: most models advertised as 'legal AI' are repackaged general-purpose models, never truly evaluated on law, compliance, or risk management tasks.

Our lab brings together machine learning researchers, legal scholars, and regulatory compliance experts based in Paris and Tunis. Every model we publish is first a research milestone: trained, tested, and benchmarked before being deployed on the Contractzlab platform and relied upon daily by banks, insurers, law firms, and enterprise corporations.

  • Applied research, validated in real-world production conditions
  • Multidisciplinary team: ML, law, regulatory compliance
  • Presence in Paris and Tunis
  • Open publications and benchmarks for the community
Our Vision

AI-assisted law, never AI-delegated law

We believe the value of legal AI is not measured by its ability to generate persuasive text, but by its capacity to reason accurately and reproducibly under human legal oversight. Our horizon: models whose reliability is proven before deployment, with every improvement being measurable, not merely perceived.

Research Axes

Three core directions structure our work

01

Specialized Legal Reasoning

Specializing language models through reinforcement learning for contractual, regulatory, and compliance reasoning — going beyond basic information retrieval.

02

Evaluation & Benchmarking

Designing rigorous evaluation protocols — independent dual-judge frameworks and statistical significance tests — to objectively measure legal reasoning quality.

03

Risk, Compliance & Geographic Coverage

Expanding our benchmarks and models to European, African, and emerging market regulatory frameworks, ensuring reliable legal AI beyond well-documented jurisdictions.

Publications

Our Published Works

Each publication documents the methodology, evaluation datasets, and results — reproducible and open for discussion.

Legal AI

ContractZLab Benchmark: Measuring legal, regulatory & risk reasoning in Europe, Africa & emerging markets

Multi-jurisdictional evaluation protocol

Legal AI

GraphRAG for Legal Intelligence: Structuring Contract Knowledge as a Graph

Graph-based retrieval-augmented generation applied to legal document understanding

Legal AI

LLM Compliance Benchmark: Evaluating Language Models on Regulatory Tasks

A dedicated benchmark protocol for regulatory compliance evaluation

Benchmarks

Measure before claiming

We build and utilize evaluation protocols designed to withstand scrutiny, not to flatter our results.

External Reference

LEXam

Reference benchmark for legal reasoning, used to evaluate Mike-7B against generalist models.

436
evaluated questions
+9,4 %
measured gain
Proprietary Benchmark

ContractZLab Benchmark

In-house protocol designed to evaluate legal, regulatory, and risk reasoning across jurisdictions in Europe, Africa, and emerging markets.

Multi-jurisdiction
coverage
Methodology

Independent Dual-Judge Evaluation

Every response is evaluated by two distinct judge models (GPT-4o, DeepSeek) to eliminate single-evaluator bias, backed by statistical significance testing.

2
independent AI judges
Our Philosophy

Five core principles guiding our research

Research applied to law carries a profound responsibility: never confuse linguistic fluency with logical accuracy.

Rigor before announcement

No result is communicated without a documented, statistically verified evaluation protocol.

Human oversight remains the rule

Our models assist legal decision-making; they never replace it.

Client data never trains our models

Model training is performed on strictly controlled datasets, completely separate from production data.

Reproducibility as a standard

Our methodologies are fully documented to be understood, scrutinized, and challenged.

Geographic coverage matters

A model reliable in France must be equally reliable in Africa and emerging markets — not just well-documented jurisdictions.

Collaborations

Research forged in the real world

Our work is shaped by direct feedback from real users and leading legal and compliance experts.

🏦

Financial Institutions

Real-world use cases in compliance, KYC, and contractual risk management.

⚖️

Law Firms

Expert validation of the quality of legal reasoning produced.

🎓

Academia

Methodological exchange on specialized language model evaluation.

🌍

Europe–Africa Ecosystem

Expanding our regulatory coverage across emerging markets.

Research Roadmap

What we are working on

A roadmap guided by the same principle: only publish what has been measured.

In Progress

Expansion of the Mike Model Family

Broadening reasoning capabilities across new contractual and regulatory domains.

Next Step

Expanded Regulatory Coverage

Strengthening the ContractZLab Benchmark across new African and emerging market jurisdictions.

Upcoming

Progressive Open Access to Benchmarks

Sharing evaluation protocols with the research community to enhance legal model comparability.

Long-term Vision

Reliability Standards for Legal AI

Helping define industry-standard evaluation criteria for language models applied to law and compliance.

Researchers, legal professionals, institutions: let's build the next generation of legal AI together

Contractzlab Research collaborates with academic teams, financial institutions, and law firms eager to contribute to our evaluation work.