My research themes lie in the broad area of heterogeneous data integration and exploitation, including heterogeneous and multi-modal data as well as warehouse, data lake and lakehouse architectures.
The general scientific questions that are driving me every day include:
How to effectively and efficiently collect, organize and store heterogeneous data produced by various actors?
How to explore and exploit large amounts of data, especially for domain experts?
How to clean, join, merge, and sementically enrich raw data for better decision making?
My research applies to various domains including sustainable cities, media, and healthcare, with a strong interest in sustainable cities.
Projects and tools
Current projects
EXPERT-MED-IA
Le confrère virtuel d'aide à la décision pour le soignant
October 2026 - October 2029
member
AURA region (2026-2028)
Scientific achievements
Practical outcomes
Associated publications
TODO
Artifacts
MOBI4ALL
Towards intelligent systems for inclusive and privacy-respecting mobility for individuals with Autism Spectrum Disorder (ASD)
October 2026 - October 2030
member (WP4)
ANR PRC (2026-2030)
Scientific achievements
Practical outcomes
Associated publications
TODO
Artifacts
Previous projects
BETTER
Real-world health-data distributed analytics research platform
heterogeneous dataE-R modelFAIR principlesclinical data
Two conceptual models for data and metadata respectively, both general, able to represent various multi-modal healthcare data, allowing the usage of ontologies, and extensible to various healthcare scenarios
An ETL algorithm to fully automatically convert existing data and metadata to instances of our models, resulting in an interoperable database
A metadata and data catalogue for exploration and ease federated learning algorithm design.
Practical outcomes
Each hospital has an interoperable database, part of the global BETTER network
Federated learning algorithms are co-created and implemented by researchers and practitioners
The catalogue implementation for the 3 use-cases (genetic rare diseases)
Efficiently enumerate all paths connecting named entities appearing in multi-modal and heterogeneous data by relying on an intermediate summary graph and a view-based algorithm
The ranking of paths based on their interestingness, a quantitative measure based on entity confidence and information dilution
Practical outcomes
A novel ChatGPT-based module for entity extraction
The efficient evaluation of paths by using a dedicated multi-query optimization algorithm
Associated publications
TODO
Artifacts
PathWays (Java application, 4k LOC, 1.8 year)
ABSTRA
Entity-Relationship summaries from semi-structured data
A novel model-agnostic data summarization method relying on data kinds instead of specific model features
The selection of most representative entities, their attributes and relationships from the data summary by relying on a graph representation and PageRank node scores
Practical outcomes
A end-to-end pipeline to summarize any (set of) heterogeneous datasets
An interface to visualize summaries as Entity-Relationship diagrams
Associated publications
TODO
Artifacts
Abstra (Java application, 10k LOC, 3.2 years)
PREDIHOOD
Supervised prediction of the environment of neighborhoods