Contractzlab Research is the AI laboratory that designs, trains, and evaluates the models driving the Contractzlab platform. Our mission: produce rigorous research in law, compliance, and risk management, and turn it into reliable systems for demanding organizations.
Contractzlab Research stems from a simple observation: most models advertised as 'legal AI' are repackaged general-purpose models, never truly evaluated on law, compliance, or risk management tasks.
Our lab brings together machine learning researchers, legal scholars, and regulatory compliance experts based in Paris and Tunis. Every model we publish is first a research milestone: trained, tested, and benchmarked before being deployed on the Contractzlab platform and relied upon daily by banks, insurers, law firms, and enterprise corporations.
We believe the value of legal AI is not measured by its ability to generate persuasive text, but by its capacity to reason accurately and reproducibly under human legal oversight. Our horizon: models whose reliability is proven before deployment, with every improvement being measurable, not merely perceived.
Specializing language models through reinforcement learning for contractual, regulatory, and compliance reasoning — going beyond basic information retrieval.
Designing rigorous evaluation protocols — independent dual-judge frameworks and statistical significance tests — to objectively measure legal reasoning quality.
Expanding our benchmarks and models to European, African, and emerging market regulatory frameworks, ensuring reliable legal AI beyond well-documented jurisdictions.
Each publication documents the methodology, evaluation datasets, and results — reproducible and open for discussion.
Multi-jurisdictional evaluation protocol
Graph-based retrieval-augmented generation applied to legal document understanding
A dedicated benchmark protocol for regulatory compliance evaluation
We build and utilize evaluation protocols designed to withstand scrutiny, not to flatter our results.
Reference benchmark for legal reasoning, used to evaluate Mike-7B against generalist models.
In-house protocol designed to evaluate legal, regulatory, and risk reasoning across jurisdictions in Europe, Africa, and emerging markets.
Every response is evaluated by two distinct judge models (GPT-4o, DeepSeek) to eliminate single-evaluator bias, backed by statistical significance testing.
Research applied to law carries a profound responsibility: never confuse linguistic fluency with logical accuracy.
No result is communicated without a documented, statistically verified evaluation protocol.
Our models assist legal decision-making; they never replace it.
Model training is performed on strictly controlled datasets, completely separate from production data.
Our methodologies are fully documented to be understood, scrutinized, and challenged.
A model reliable in France must be equally reliable in Africa and emerging markets — not just well-documented jurisdictions.
Our work is shaped by direct feedback from real users and leading legal and compliance experts.
Real-world use cases in compliance, KYC, and contractual risk management.
Expert validation of the quality of legal reasoning produced.
Methodological exchange on specialized language model evaluation.
Expanding our regulatory coverage across emerging markets.
A roadmap guided by the same principle: only publish what has been measured.
Broadening reasoning capabilities across new contractual and regulatory domains.
Strengthening the ContractZLab Benchmark across new African and emerging market jurisdictions.
Sharing evaluation protocols with the research community to enhance legal model comparability.
Helping define industry-standard evaluation criteria for language models applied to law and compliance.
Contractzlab Research collaborates with academic teams, financial institutions, and law firms eager to contribute to our evaluation work.