SynDAiTE: Synthetic Data for AI Trustworthiness and Evolution
Workshop at the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML-PKDD 2026), September 7, 2026 - Naples, Itay
Organisers
University of Camerino | Vici & C.
University of Porto | Fraunhofer AICOS
Brunel University London | MHRA UK
🚨📢 Where to submit? 📢🚨
All submissions must be done via CMT.
Table of contents
Aims and Scope
The rapid advancement of artificial intelligence (AI) relies heavily on access to large, diverse, and high-quality datasets for training and evaluation. However, the increasing scarcity of data, strict privacy regulations, and the high costs associated with collection and annotation are creating significant barriers to progress. Projections suggest that by 2050, we may face a shortage of fresh text data, and by 2060, image data may become similarly limited. These challenges make it imperative to explore alternatives that can sustain AI’s growth and effectiveness. Synthetic data offers a compelling solution to these issues, with the advantages of scalability, customisation, and inherent anonymisation. It allows for the generation of large volumes of tailored datasets without the same privacy and cost concerns of real data.
Important Dates
- Paper Submissions: June 19th, 2026
- Notifications: July 20th, 2026
- Camera-Ready: July 30th, 2026
- Workshop: September 11th, 2026
All deadlines are 11:59 pm, Pacific Time.
Topics
SynDAiTE welcomes contributions on the use of synthetic data on all topics below, independent of the application domain (e.g., health, finance, business, basic sciences, construction, computational advertising, IoT, etc.) and of data types (e.g., networks, graphs, logs, spatiotemporal, multimedia, time series, genomic sequences, and streaming data.):
- Generation:
- Techniques for high-fidelity, domain-specific synthetic data production.
- Customisation for anomaly and rare-event detection.
- Scalability and adaptability to various applications.
- Ethical aspects of synthetic data.
- Responsible AI:
- Meta-learning for understanding models and algorithms.
- Privacy-preserving ML.
- Methodologies to support evaluation according to Responsible AI pillars.
- AI Auditing and Red Teaming:
- Stress testing models and algorithms.
- Interpretability and explainability for auditing ML systems.
- Adversarial ML.
- Challenges in Synthetic Data Use:
- Fidelity and accuracy concerns in real-world applications.
- Bias detection and mitigation strategies.
- Validation frameworks to ensure reliability and generalisation.
- Dynamic and Temporal Contexts:
- Generating data for streaming environments and micro-batch processing.
- Incorporating temporal complexity and drift phenomena.
- Learning Frameworks:
- Online continual learning with synthetic datasets.
- Data stream mining with synthetic drifts and anomalies.
- Supervised and unsupervised learning with synthetic augmentation.
- Evaluating synthetic data’s impact on model performance and robustness.
- Anomaly, Novelty, and Drift Detection:
- Leveraging synthetic data for rare-event detection in evolving datasets.
- Mitigating challenges associated with concept drift and changing data distributions.
- Using synthetic data to test and refine anomaly detection frameworks.
- Simulating anomalies in stream/dynamic environments.
- Applications of Data Synthesis for AI training (Healthcare, Finance, Transportation, Cybersecurity).
- Explainable Synthetic Data Generation:
- Explainable synthesis methods, including intrinsically interpretable generators (e.g., structured/constrained models, probabilistic graphical models) and explainability-enhanced approaches.
- Frameworks and tools for tracing, auditing, and evaluating explainability in synthetic data Generation.
- LLM- and Agent-Based Synthetic Data Generation:
- Synthetic data generation using large language models and foundation models.
- Agent-based approaches for synthetic data generation.
- Hybrid methods combining LLMs with probabilistic, causal, or structured models.
- Evaluation of reliability, bias, controllability, and explainability in LLM-generated synthetic.
Invited Talks:
The Interplay of Missing Data and Synthetic Data in Machine Learning - Ricardo Cardoso Pereira
- 12:20
|
Abstract:
Missing data refers to values we failed to collect, while synthetic data deals with values we generate on purpose. Consequently, the two fields are closely related. Imputation can be seen as conditional generation, and generative models are very often used to fill in incomplete data. In the other direction, missingness can impact synthetic data: generators trained on incomplete or badly imputed data can inherit their biases, while artificially injected missingness is useful for benchmarking and privacy. This talk covers both directions of this relationship, including open challenges.
Speaker's Bio:
Ricardo Cardoso Pereira is an Assistant Professor at the Department of Informatics Engineering at the University of Coimbra, where he is also affiliated with the Centre for Informatics and Systems (CISUC). His academic and research work is focused on artificial intelligence, with specific interests in machine learning, deep learning, data-centric AI, data quality (particularly addressing missing data), and fairness in machine learning. He holds a PhD in Artificial Intelligence and an MSc in Computer Science from the University of Coimbra.
|
Program at a Glance (Room TBD):
- 10:35 - Welcome and Opening Remarks
- 10:45:45 - Short Talks (Part A)
- 10:45 -
- Ibrahim et al. - Demographically-Conditioned Synthetic Medical Images for Bias Mitigation and Bias Detection in Disease Classifiers
- 10:53 -
- Burlova et al. - MechFaithBench: A Counterfactual Benchmark for Visual Grounding Diagnostics in Vision-Language Models
- 11:01 -
- Kapenekakis et al. - Amalgam: Hybrid LLM-PGM Synthesis Algorithm for Accuracy and Realism
- 11:09 -
- Komorniczak et al. - How well does Classification Accuracy capture Concept Drift Detection Quality? An overview of Concept Drift Detection evaluation
- 11:20 - Poster Session (Part A)
- 11:20 - Oral Presentations (Part B)
- 11:20 -
- Youssef et al. - PATE-TabTransGAN: Differentially Private Synthetic Tabular Data Generation via Transformer-Based Student Discrimination
- 11:28 -
- Atmika et al. - Train, Test, Re-evaluate: Schedule-Sensitive Evaluation of Generative Data for Hand Detection
- 11:36 -
- Kokkinogenis et al. - Berry Picking Across Data Landscapes: Understanding Performance Sensitivity in Collaborative Filtering
- 11:44 -
- Oliveira et al. - Data Complexity Effects on Synthetic Data Quality, Privacy, and Utility
- 11:55 - Poster Sessions (Part B)
- 12:20 - - Ricardo Cardoso Pereira - The Interplay of Missing Data and Synthetic Data in Machine Learning
- 12:50 - Concluding remarks
Registration and Presentation Policy
Each accepted paper must have at least one author registered for the full conference by the early registration deadline and must be presented at the workshop even if they opt-out of the post-proceedings. We expect the authors, the program committee, and the organizing committee to adhere to the ECML-PKDD Code of Conduct.
The Main Conference organization team will manage the registration: https://ecmlpkdd.org/2026/attending-registration/
Acknowledgement
“The Synthetic Data for AI Trustworthiness and Evolution (SynDAiTE 2026)” workshop has been supported by VICI & C.
For general inquiries about the workshop, please email syndaite@gmail.com