We bridge machine learning, data systems, and discovery ...
Our research connects machine learning and data engineering to make AI models useful under real-world constraints on resources, data, and computing systems. We design learning methods by considering data access, computing architecture, and available resources together from the outset. This perspective spans efficient access to very large data collections, scalable execution on high-performance systems, and compact models for edge and TinyML hardware. Current funded projects include AI4Forest and TinyAIoT.
Large-Scale Machine Learning
Many large-scale applications are limited less by a model's arithmetic than by data access and movement. We therefore couple learning methods with index structures, storage systems, and execution pipelines.
In search-by-classification, a small number of examples describe the objects a user wants to retrieve. Our index-aware decision trees translate inference into range queries, enabling interactive searches across billions of objects. RapidEarth applies this idea to satellite-image archives, while CLIP-Branches combines it with interactive fine-tuning of multimodal models.
- Christian Lülf, Denis Mayr Lima Martins, Salles Marcos Antonio Vaz, Yongluan Zhou & Fabian Gieseke (2024). CLIP-Branches: Interactive Fine-Tuning for Text-Image Retrieval. Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, Demo Track. SIGIR 2024.
- Christian Lülf, Denis Mayr Lima Martins, Salles Marcos Antonio Vaz, Yongluan Zhou & Fabian Gieseke (2023). RapidEarth: A Search Engine for Large-Scale Geospatial Imagery. Proceedings of the 31st International Conference on Advances in Geographic Information Systems, Demo Paper. SIGSPATIAL 2023.
- Christian Lülf, Denis Mayr Lima Martins, Salles Marcos Antonio Vaz, Yongluan Zhou & Fabian Gieseke (2023). Fast Search-By-Classification for Large-Scale Databases Using Index-Aware Decision Trees and Random Forests. Proceedings of the VLDB Endowment, 16, 2845–2857. VLDB 2023.
Current research projects often process hundreds of terabytes of satellite observations. Here, transferring the data can become as important a bottleneck as computation. We develop trainable selection masks that identify and transfer only the parts of an input that matter for a task. For petabyte-scale archives, automated preselection can substantially reduce data movement and end-to-end inference time.
This line of work directly supports the scalable Earth-observation pipelines developed in AI4Forest.
- Jan Pauls, Max Zimmer, Una M. Kelly, Martin Schwartz, Sassan Saatchi, Philippe Ciais, Sebastian Pokutta, Martin Brandt & Fabian Gieseke (2024). Estimating Canopy Height at Scale. 41st International Conference on Machine Learning. ICML 2024.
- Stefan Oehmcke & Fabian Gieseke (2022). Input Selection for Bandwidth-Limited Neural Network Inference. Proceedings of the 2022 SIAM International Conference on Data Mining, 280–288. SDM 2022.
Tiny Machine Learning
TinyML brings training and inference to severely resource-constrained devices. Adapted training procedures and memory layouts reduce the resource requirements of boosted-tree models by factors of 4–16 while preserving predictive performance. Trainable quantization further reduces the sensor data that must be transmitted. Applications include privacy-preserving bicycle counting and energy-efficient bird-species recognition.
These methods are developed and evaluated in the TinyAIoT project and related environmental-monitoring initiatives such as Birdiary.
- Nina Herrmann, Jan Stenkamp, Benjamin Karic, Stefan Oehmcke & Fabian Gieseke (2026). Boosted Trees on a Diet: Compact Models for Resource-Constrained Devices. The Fourteenth International Conference on Learning Representations. ICLR 2026.
- Karsten Schrödter, Jan Stenkamp, Nina Herrmann & Fabian Gieseke (2026). Trainable Bitwise Soft Quantization for Input Feature Compression. Third Conference on Parsimony and Learning. CPAL 2026.
- Anni Henriikka Kurkela, Jan Stenkamp, Paula Scharf, Thomas Bartoschek & Fabian Gieseke (2026). TinyML for Environmental Monitoring: Bird Species Image Classification on Resource-Constrained Devices. European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, Applied Data Science Track. ECML PKDD 2026.
Machine Learning and High-Performance Computing
Our work exploits the parallelism of modern many-core systems through hardware-aware algorithm design. Buffer k-d trees batch search requests for massively parallel nearest-neighbor search on GPUs. Instead of traversing a tree independently for every query, buffers collect queries at tree nodes and process them together. This reorganizes irregular control flow into large, hardware-friendly batches and improves memory access on GPUs.
We have extended this perspective to data-intensive learning models and scientific pipelines, including parallel regression for satellite time series. In change detection, such methods can reduce computations over billions of time series from weeks or years to hours or days. The central principle is to co-design algorithms, memory movement, vectorization, and distributed execution for the target architecture.
- Fabian Gieseke, Sabina Rosca, Troels Henriksen, Jan Verbesselt & Cosmin Eugen Oancea (2020). Massively-Parallel Change Detection for Satellite Time Series Data with Missing Values. Proceedings of the 36th IEEE International Conference on Data Engineering, 385–396. ICDE 2020.
- Fabian Gieseke & Christian Igel (2018). Training Big Random Forests with Little Resources. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 1445–1454. KDD 2018.
- Fabian Gieseke, Justin Heinermann, Cosmin E. Oancea & Christian Igel (2014). Buffer k-d Trees: Processing Massive Nearest Neighbor Queries on GPUs. Proceedings of the 31st International Conference on Machine Learning, 172–180. ICML 2014.
Applications
Alongside our methodological research, we develop machine-learning techniques with experts from Earth observation, astrophysics, and other application domains.
Earth Observation
Our work ranges from detecting abrupt ecosystem changes and mapping individual trees and their carbon stocks to creating global, time-dependent canopy-height maps. We build efficient inference pipelines for petabyte-scale collections, develop foundation representations for environmental monitoring, and study predictive uncertainty through quantile regression.
Much of this work is carried out within AI4Forest; interactive results and further material are also available at ai4forest.eu.
- Karsten Schrödter, Jan Pauls & Fabian Gieseke (2026). Canopy Tree Height Estimation using Quantile Regression: Modeling and Evaluating Uncertainty in Remote Sensing. Twenty-Ninth Annual Conference on Artificial Intelligence and Statistics. AISTATS 2026.
- Ibrahim Fayad, Max Zimmer, Martin Schwartz, Philippe Ciais, Fabian Gieseke, Gabriel Belouze, Sarah Brood, Aurelien De Truchis & Alexandre d'Aspremont (2025). DUNIA: Pixel-Sized Embeddings via Cross-Modal Alignment for Earth Observation Applications. 42nd International Conference on Machine Learning. ICML 2025.
- Jan Pauls, Max Zimmer, Berkant Turan, Sassan Saatchi, Philippe Ciais, Sebastian Pokutta & Fabian Gieseke (2025). Capturing Temporal Dynamics in Large-Scale Canopy Tree Height Estimation. 42nd International Conference on Machine Learning. ICML 2025.
- Paulo Negri Bernardino, Wanda De Keersmaecker, Stéphanie Horion, Ruben Van De Kerchove, Stef Lhermitte, Rasmus Fensholt, Stefan Oehmcke, Fabian Gieseke, Koenraad Van Meerbeek, Christin Abel, Jan Verbesselt & Ben Somers (2025). Predictability of Abrupt Shifts in Dryland Ecosystem Functioning. Nature Climate Change, 15, 86–91.
- Maurice Mugabowindekwe, Martin Brandt, Jerome Chave, Florian Reiner, David Skole, Ankit Kariryaa, Christian Igel, Pierre Hiernaux, Philippe Ciais, Ole Mertz, Xiaoye Tong, Sizhuo Li, Gaspard Rwanyiziri, Thaulin Dushimiyimana, Alain Ndoli, Uwizeyimana Valens, Jens-Peter Lillesø, Fabian Gieseke, Compton Tucker, Sassan S Saatchi & Rasmus Fensholt (2022). Nation-wide mapping of tree-level aboveground carbon stocks in Rwanda. Nature Climate Change.
- Martin Brandt, Compton J. Tucker, Ankit Kariryaa, Kjeld Rasmussen, Christin Abel, Jennifer Small, Jerome Chave, Laura Vang Rasmussen, Pierre Hiernaux, Abdoul Aziz Diouf, Laurent Kergoat, Ole Mertz, Christian Igel, Fabian Gieseke, Johannes Schöning, Sizhuo Li, Katherine Melocik, Jesse Meyer, Scott Sinno, Eric Romero, Erin Glennie, Amandine Montagu, Morgane Dendoncker & Rasmus Fensholt (2020). An unexpectedly large count of trees in the West African Sahara and Sahel. Nature.
Astrophysics
Astronomy and particle physics were important application areas in an earlier phase of our research. Modern sky surveys produce very large image and catalogue collections in which rare, scientifically relevant objects must be identified among millions or billions of observations. Manual inspection is therefore impossible, making accurate and scalable machine-learning pipelines essential.
For astronomical surveys, we developed nearest-neighbor methods for photometric redshift estimation and for discovering previously unknown high-redshift quasars. We also designed convolutional neural networks that distinguish genuine transient events from imaging artefacts and thereby reduce the number of candidates requiring expert review. These projects established our focus on scalable inference, efficient candidate selection, and close collaboration with domain scientists.
- Fabian Gieseke, Steven Bloemen, Cas van den Bogaard, Tom Heskes, Jonas Kindler, Richard A. Scalzo, Valerio A.R.M. Ribeiro, Jan van Roestel, Paul J. Groot, Fang Yuan, Anais Möller & Brad E. Tucker (2017). Convolutional Neural Networks for Transient Candidate Vetting in Large-Scale Surveys. Monthly Notices of the Royal Astronomical Society, 472, 3101–3114. MNRAS.
- Kai Lars Polsterer, Peter Zinn & Fabian Gieseke (2013). Finding New High-Redshift Quasars by Asking the Neighbours. Monthly Notices of the Royal Astronomical Society, 428, 226–235. MNRAS.
Smart Cities and Smart Grids
Smart cities and smart grids combine distributed sensing, local intelligence, and networked infrastructure. Our research addresses forecasting for increasingly dynamic energy systems: the growing share of weather-dependent renewable generation makes energy supply more volatile, while electricity demand varies across time, locations, and consumer groups. We use state-of-the-art machine-learning techniques to obtain accurate forecasts of renewable-energy production and energy demand from historical measurements, weather information, calendar effects, and other contextual signals. These forecasts can support grid operation, flexibility planning, and the reliable coordination of generation, storage, and consumption.
For distributed smart-city sensing systems, we develop compact models that process data close to where it is collected and transmit only the results. This reduces bandwidth and energy consumption while supporting privacy-aware monitoring and timely decisions. One application is the monitoring of bicycle-parking facilities: a camera-equipped ESP32-S3 microcontroller runs a compressed object-detection model locally and sends only the resulting bicycle count via LoRaWAN to a remote service.