Skip to content
Research

Bridging the GAP between
Computer Vision and Generative AI

We work on problems we find genuinely interesting in Computer Vision and Generative AI for both generation and perception and their intersection.

Research Areas

Diffusion Models & Generative Synthesis

Diffusion Models & Generative Synthesis

Controllable image and video generation with diffusion models — from fine-grained, part-level editing to consistent character animation and 3D-aware scene layout.

Multi-Modal LLMs & Agentic Systems

Multi-Modal LLMs & Agentic Systems

Vision-language models and agentic pipelines that reason over images, diagrams and video — not just generate text — and act on that reasoning.

Vision Representation Learning

Vision Representation Learning

Learning visual embeddings that capture what actually matters — identity, semantics, and editing intent — instead of entangling it with background context or surface appearance.

3D Vision & Scene Understanding

3D Vision & Scene Understanding

Shape correspondence, keypoint reasoning, and language-guided object placement in real 3D scenes — connecting geometric understanding with natural-language interaction.

Uncertainty-Aware Perception

Uncertainty-Aware Perception

Confidence-propagating networks for sparse and noisy data — depth completion, optical flow, and regression tasks where knowing what the model doesn't know matters.

Detection, Tracking & Video Segmentation

Detection, Tracking & Video Segmentation

Robust object detection, tracking and segmentation in real video — including distractor-aware tracking and zero-shot segmentation built on pre-trained diffusion models.

Publications

2026

FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations
ICML 2026

FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations

Fedor Rodionov, Abdelrahman Eldesokey, Michael Birsak, John Femiani, Bernard Ghanem, Peter Wonka

A diagnostic benchmark for spatial reasoning in LLMs, built on structured JSON/XML representations of indoor scenes and covering tasks like distanc…

DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation
ECCV 2026

DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation

Qian Wang, Zhenyu Li, Abdelrahman Eldesokey, Peter Wonka

A post-training framework for subject-driven image generation that optimizes structural diversity and identity consistency together, treating ident…

CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models
Preprint 2026

CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models

Reem Alzahrani, Hassan Alshanqiti, Bushra Bin Hemid, Zaid Alyafeai, Abdelrahman Eldesokey, Bernard Ghanem

Tests whether vision-language models actually look at the image when counting objects, using paired factual/counterfactual images with edited count…

Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation
Preprint 2026

Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation

Abdelrahman Eldesokey, Merey Ramazanova, Ahmad Sait, Ansar Khangeldin, Karen Sanchez, Tong Zhang, Bernard Ghanem

Proposes matching each text-to-image evaluation skill (e.g. counting, spatial relations, attribute binding) to an annotation strategy suited to its…

SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy
Preprint 2026

SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy

Ismael Elsharkawi, Ahmed Sait, Silvio Giancola, Bernard Ghanem, Hossam Sharara, Abdelrahman Eldesokey

Introduces a benchmark for whether soccer-video VLMs are actually grounded in what's on screen, not just pattern-matching to the right event label.…

NearID: Identity Representation Learning via Near-Identity Distractors
ECCV 2026

NearID: Identity Representation Learning via Near-Identity Distractors

Aleksandar Cvejic, Rameen Abdal, Abdelrahman Eldesokey, Bernard Ghanem, Peter Wonka

Trains identity-aware visual representations by contrasting each identity against visually similar but distinct "near-identity" instances, separati…

2025

Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation
NeurIPS 2025Spotlight

Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation

Abdelrahman Eldesokey, Aleksandar Cvejic, Bernard Ghanem, Peter Wonka

Introduces a method for detecting visual inconsistencies in subject-driven image generation by leveraging visual correspondence, improving reliabil…

EditCLIP: Representation Learning for Image Editing
ICCV 2025

EditCLIP: Representation Learning for Image Editing

Qian Wang, Aleksandar Cvejic, Abdelrahman Eldesokey, Peter Wonka

A representation-learning approach tailored to image editing, learning embeddings that capture the semantics of an edit itself rather than just ima…

PlaceIt3D: Language-Guided Object Placement in Real 3D Scenes
ICCV 2025

PlaceIt3D: Language-Guided Object Placement in Real 3D Scenes

Ahmed Abdelreheem, Filippo Aleotti, Jamie Watson, Zawar Qureshi, Abdelrahman Eldesokey, Peter Wonka, Gabriel Brostow, Sara Vicente, Guillermo Garcia-Hernando

Enables placing objects into real 3D scenes using natural-language instructions, bridging language understanding with geometric scene reasoning.

ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models
ICCV 2025

ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models

Bingchen Gong, Diego Gomez, Abdullah Hamdi, Abdelrahman Eldesokey, Ahmed Abdelreheem, Peter Wonka, Maks Ovsjanikov

Shows that large language models can be used for zero-shot, point-level reasoning to detect semantic 3D keypoints without task-specific training da…

PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion Models
SIGGRAPH 2025

PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion Models

Aleksandar Cvejic*, Abdelrahman Eldesokey*, Peter Wonka

A method for precise, part-level image editing built on pre-trained diffusion models, allowing targeted edits to specific object parts without dist…

VidSeg: Zero-Shot Video Semantic Segmentation based on Pre-Trained Diffusion Models
CVPR 2025

VidSeg: Zero-Shot Video Semantic Segmentation based on Pre-Trained Diffusion Models

Qian Wang, Abdelrahman Eldesokey, Mohit Mendiratta, Fangneng Zhan, Adam Kortylewski, Christian Theobalt, Peter Wonka

Repurposes pre-trained diffusion models for zero-shot semantic segmentation of video, removing the need for task-specific labeled training data.

Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image Generation
ICLR 2025

Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image Generation

Abdelrahman Eldesokey, Peter Wonka

Gives users interactive, explicit control over 3D object layout when generating images with diffusion models, closing the gap between free-form pro…