Below you find the tentative program of the conference.
Tuesday, September 22
12:30 - 15:00 Women in Vision Workshop at GCPR 2026
Organizers: Dr. Natacha Kuete Meli (University of Siegen), Prof. Michael Moeller (University of Siegen) The Women in Vision Workshop at GCPR 2026 aims to celebrate the impact of women in the field, promote visibility, and foster inclusiveness and diversity in pattern recognition and related research areas. The workshop provides an opportunity for early-career researchers and talented women scientists to connect with one another. Finger food and coffee will be provided to enrich the discussions.
Date: September 22, 2026 - 12:30–14:30 - 12:30–12:45 Welcome by Dr. Joana Grah - 12:45–13:10 Invited talk by Prof. Zorah Lähner - 13:10–13:35 Invited talk by Dr. Rachel Hegemann - 13:35–13:45 Preparation for the mentoring session - 13:45–14:30 Mentoring session with mentors Dr. Natacha Kuete Meli, Dr. Joana Grah, Dr. Rachel Hegemann, Prof. Zorah Lähner, and more to come. - 14:10–15:00 Poster session
About the Speakers Dr. Joana Grah (Robert Koch Institute, joanasarahgrah.science) is a Research Associate in applied mathematics at the Robert Koch Institute. She holds a PhD in Applied Mathematics from the University of Cambridge, UK. Beyond her research, she actively advocates for equal opportunities and feminism to support science in diverse, interdisciplinary, and empathetic environments. She is a co-founder of “Her Math Story,” (https://joanasarahgrah.science/index.php/her-maths-story/) a platform showcasing the stories of women mathematicians across a wide range of careers, non-linear paths, and individual decision-making processes. Prof. Zorah Lähner (University of Bonn and the Lamarr Institute, geometryinml.cs.uni-bonn.de/team/zorah/) has been a full professor at the University of Bonn and the Lamarr Institute since 2026, where she leads the Geometry in Machine Learning group. She obtained her PhD in Computer Science from the Technical University of Munich. Her research focuses on geometric aspects of optimization, including geometric deep learning and 3D geometry processing. Dr. Rachel Hegemann (Deutsche Bahn AG) is a Data Scientist at Deutsche Bahn AG. She holds a PhD in spatio-temporal social networks and data reconstruction from the University of California, Los Angeles. Her current work focuses on developing assessment and testing strategies to ensure AI safety and quality in industrial applications.
Vision foundation models trained on massive, proprietary datasets currently dominate the field, creating a significant barrier for the broader research community. This tutorial shifts the narrative, demonstrating how researchers can train vision encoders entirely on open-source datasets to match or even surpass the performance of closed-source giants. We move beyond theoretical assumptions to break down the exact engineering requirements needed to close this gap whether you are working with a massive cluster or a multi-node setup.
Scaling models introduce severe optimization challenges, from loss spikes to representation collapse. We detail how to maintain training stability when moving to a few hundred million parameter model, but we also translate these lessons into practical strategies for smaller-scale training. By monitoring gradients and architectural bottlenecks, researchers can ensure steady convergence and faster training times, regardless of their total FLOPs budget. The most significant lever for performance isn't just more compute, it’s better data. Moving away from indiscriminate web scraping, we examine how advanced filtering strategies such as semantic deduplication and entropy-based sampling directly dictate a model’s zero-shot transfer capabilities. These techniques allow researchers to achieve superior feature quality with a fraction of the data, making high-performance training accessible to those without petascale storage. What this tutorial aims to achieve:
Demonstrate how vision encoders trained exclusively on public data (like DataComp or LAION) can achieve parity with proprietary models across standard benchmarks.
Provide architectural and hyperparameter strategies to maintain steady convergence when scaling vision foundation models.
Teach algorithmic data filtering techniques that shift the focus from raw data volume to high-quality, balanced datasets, effectively improving model robustness.
Bridge the gap between "big lab" infrastructure and academic research, providing a roadmap for resource-efficient foundation model development.
15:30 - 18:30 Tutorial on Topological Data Analysis on Surface Meshes (US-C 115)
Organizers: Jonas Lukasczyk Target Audience: Novice Workshop session: Half day
This tutorial introduces the foundations and practical applications of Topological Data Analysis (TDA) for surface meshes. TDA has emerged as a powerful framework for extracting robust, multi-scale structural features from complex data, making it particularly valuable for visualization, geometry processing, and scientific computing.
In this tutorial, participants will learn how to use the Topology Toolkit (TTK, https://topology-tool-kit.github.io/), a software library that provides a wide range of TDA algorithms for surface meshes, such as topological skeletonization and remeshing, the computation of persistent generators for detecting connectivity, loops, and holes, as well as topology-driven shape matching techniques. TTK is integrated into ParaView, which can be easily installed on Linux, Windows, and MacOS (https://www.paraview.org/download/). The tutorial is designed to be highly interactive, where attendees will engage in guided, hands-on exercises that teach them how to apply TDA methods within their own workflows.
Participants are asked to kindly bring their laptops and install paraview with the TTK-plugin in advance of the tutorial.
Beyond Pixels and Points: Explicit Control and Implicit Physics for 4D Generation
Abstract: The creation and understanding of 4D dynamics remain a key challenge spanning computer vision and graphics. However, modeling the temporal evolution of the 3D world requires balancing two competing factors: explicit representations that offer fine-grained control and structural fidelity, and implicit (neural) models that capture scalable and complex visual priors.
In this talk, I will discuss our efforts to bridge explicit and implicit paradigms to reconstruct, generate, and simulate 4D content. First, we will examine LooseControlVideo, demonstrating how we can condition large-scale implicit video generation models on rough, explicit 3D bounding boxes to achieve highly controllable dynamic scenes. Moving from generation to reconstruction, I will showcase how we can lift video priors into explicit, animated, and fully articulated 3D character meshes, capturing complex motion without multi-view setups. Finally, we will look beyond appearance toward underlying physical laws with Neural Voxel Dynamics. By building an implicit model of physics on top of V-JEPA feature spaces, I will describe the possibility of learning a generalized neural simulator. Together, these approaches sketch a possible trajectory towards 4D assets necessary to build fully controllable, physically grounded dynamic worlds.
Bio: Niloy J. Mitra leads the Smart Geometry Processing group in the Department of Computer Science at University College London and the Adobe Research London Lab. He received his Ph.D. from Stanford University under the guidance of Leonidas Guibas. His research develops machine learning frameworks for reconstructing/generating high-quality geometric and dynamic content in computer graphics applications. He has received several recognitions, including the ACM SIGGRAPH Significant New Researcher Award (2013), the BCS Roger Needham Award (2015), and the Eurographics Outstanding Technical Contributions Award (2019). He was elected a Eurographics Fellow in 2021, served as Technical Papers Chair for SIGGRAPH in 2022, and was inducted into the SIGGRAPH Academy in 2023. Beyond research, Niloy is an avid DIYer and enjoys reading, cricket, and cooking. More information is available at geometry.cs.ucl.ac.uk.
2D versus 2.5D Layer Arrangement for Animal Behaviour Multiplex Networks
Stefan Paul Feyer, Karsten Klein, Stephen Kobourov, Katherine Snell, Natalia Borrego, Genevieve Erin Finerty, Rob S.A. van Bemmelen, Nina J O'Hanlon, Uakendisa Muzuma, Falk Schreiber
Design and Evaluation of Fractal Shapes for Hierarchy Node Identification
Tobias Mertz, Steven Lamarr Reynolds-Ringer, Maria Borchert, Lasse Zimmer, Jörn Kohlhammer
DEXPRO – Visual Data Explorer for Production and Design Data
Florian Steinwidder, Julian Rakuschek
Parallel vs. Radial Axis Layouts in Parallel Coordinates
Gabriel Borrelli, Lars Linsen
16:00 - 17:00 Poster Session 2 (Foyer + US-C 102)
ID
Paper Title
Authors
P2-1
Pairwise Post-Hoc Cross-View Refinement of Monocular Metric 3D Geometry
Ulas Gunes; Matias Turkulainen; Mikhail Silaev; Juho Kannala; Esa Rahtu
P2-2
Ordered Diffusion for 3D Human Registration
Mattia Masiero; Ilya Petrov; Daniel Cremers; Gerard Pons-Moll; Riccardo Marin
P2-3
Diffusion-Based SAR Training Data Synthesis Controlled by Spatial Annotations
Sylvia Hochstuhl; Horst Hammer; Antje Thiele; Tobias Brosch; Padraig Davidson; Tim Remiger; Michael Teutsch
P2-4
STIP+: Context-Aware Human–Object Interaction Detection with Adaptive Relation Sampling
Yanxi Lin; Noha Sarhan; Simone Frintrop
P2-5
SDFFormer: Robust Continuous Surface Extraction from Sparse Unposed Images via Learned Geometric Priors
Adrien Schockaert; Hazem Wannous; Guillaume Dufaye; Vincent Magnier; Jean-François Witz