The problem
Spatial Augmented Reality projects digital content directly onto physical scenes. Visual fidelity is only part of an intelligent system: it must also choose suitable content and understand which visible information belongs to the projection and which belongs to the real scene. Surface texture and projector brightness constrain generation, while their visual mixture creates ambiguity for scene interpretation.
From surface-aware generation to semantic understanding
-
1. Generate content for the surface
Turn a language prompt into projected stylization while accounting for surface texture and projector limits.
-
2. Understand projection and scene separately
Separate physical and projected layers, then describe each with projection-aware captioning.
Generating a style the surface can display
LAPIG takes a language prompt and generates a projector image that changes the style of a textured surface. Simply generating a desired image and compensating it can leave saturation and texture artifacts because the projector cannot produce every requested appearance.
Projection surface adaptation addresses this constraint during generation. Learned compensation and project-and-capture models make it possible to optimize the image without physically projecting every candidate. Content and saturation losses guide the result toward the requested style while reducing surface-related artifacts. The generated image is then projected to produce the surface stylization.
The system therefore couples semantic intent with the physical display process: the surface and projector influence what should be generated, rather than appearing only as a final correction step.
Distinguishing the projection from the physical scene
ProCap studies the complementary interpretation problem. A standard vision-language model may describe projected content as if it were part of the physical scene. ProCap separates virtual and physical layers through automated segmentation, then uses region-aware retrieval to reduce ambiguous semantic context caused by projection distortion.
The framework produces descriptions of the physical scene and projected content separately. Its RGBP benchmark and dual-captioning evaluation support this distinction explicitly, with annotations for the two layers. The source code, models, and dataset are linked in the resources below.
System and demonstrations
The two systems address different sides of the interaction loop. LAPIG connects a user prompt to surface-aware image generation and physical projection. ProCap connects a captured augmented scene to separated physical and projection descriptions. They share the need to account for the real surface and the digital layer, while providing different generation and understanding pipelines.
The demonstration above introduces both tasks. The individual project websites provide the detailed generation examples, semantic comparisons, and experimental results.
Papers and implementations
These papers document the methods and developments described above. Their authors, publication details, and available resources are listed together here.
LAPIG: Language Guided Projector Image Generation with Surface Adaptation and Stylization
Yuchen Deng, Haibin Ling, Bingyao Huang
IEEE Transactions on Visualization and Computer Graphics (TVCG), 2025
Also in IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR), 2025
Paper Project Code
ProCap: Projection-Aware Captioning for Spatial Augmented Reality
Zimo Cao, Yuchen Deng, Haibin Ling, Bingyao Huang
IEEE Conference on Virtual Reality and 3D User Interfaces (IEEE VR), 2026
Paper Project Code
Datasets
Related work
Projector Compensation provides methods for controlling captured appearance on real surfaces. Differentiable & 3D ProCam Systems supplies forward models of projection, surface geometry, and light transport. These physical foundations support intelligent SAR, where the goals also involve language, content generation, and semantic distinction.
Related research topics and projects
Research topics: ProCam & SAR