TEAM NAVER Reaffirms Its Global Vision AI Capabilities with Multiple Papers Accepted at a Leading International Conference
TEAM NAVER Reaffirms Its Global Vision AI Capabilities with Multiple Papers Accepted at a Leading International Conference
TEAM NAVER Reaffirms Its Global Vision AI Capabilities with Multiple Papers Accepted at a Leading International Conference
- 23 TEAM NAVER papers accepted at ECCV 2026, continuing a record of double-digit annual acceptances across the three major computer vision conferences
- Advances in 3D spatial reconstruction include a faster, more accurate successor to DUSt3R and technology for recognizing previously unseen objects in autonomous driving
- Four papers, including research on 360-degree spatial reconstruction and the Seoul World Model, selected as ECCV Spotlight papers
August 28, 2026
Twenty-three papers by TEAM NAVER have been accepted at ECCV 2026 (European Conference on Computer Vision), one of the world’s premier computer vision conferences, once again demonstrating its global competitiveness in spatial intelligence.
Now in its 19th edition, ECCV is widely regarded as one of the world’s most prestigious AI conferences dedicated to computer vision, covering fields such as image and video processing. The conference is held every two years, and ECCV 2026 will take place in Sweden from September 8 to 12. Technology teams from NAVER LABS, NAVER LABS Europe (NLE), NAVER AI Lab, and NAVER Cloud will present their accepted papers at the conference.
The main conference accepted 23 papers from TEAM NAVER across a broad range of vision technologies, including 3D spatial perception and reconstruction, image and video learning models, and robot navigation simulation. The large number of accepted papers on 3D spatial reconstruction in particular further validates TEAM NAVER’s technological leadership in vision AI.
First, NAVER LABS Europe introduced BLASt3R, a follow-up to DUSt3R*, which had drawn considerable attention from the global 3D vision community. (Paper: BLASt3R: bundle adjustment of any image set with multi-view matching and monocular priors) BLASt3R integrates simultaneous localization and mapping (SLAM), a technique that allows a robot to map its surroundings in real time as it moves. This enables on-device 3D reconstruction while substantially improving both speed and accuracy over its predecessor.
* Released by NAVER LABS Europe in 2023, DUSt3R is an AI tool that can reconstruct an entire space in 3D from just one or two photographs.
NAVER LABS also presented research on indoor 3D scene generation that can reconstruct a room’s layout and objects from a single 360-degree panoramic image, drawing attention from the academic community. (Paper: InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360° Image) NAVER LABS also introduced a technology that uses a vision-language model (VLM) in autonomous driving to recognize objects outside the model’s training set and place them accurately on a bird’s-eye-view (BEV) map in real time. (Paper: Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints)
The Seoul World Model, developed by NAVER AI Lab through an industry-academic collaboration with KAIST, was also accepted as a main-conference paper in recognition of its potential applications in map-based virtual exploration, digital twins, autonomous driving, and robotics simulation. (Paper: Seoul World Model: Grounding World Simulation Models in a Real-World Metropolis) The model uses NAVER’s vast spatial datasets, including panoramic imagery from across the city, to generate long-horizon videos depicting Seoul as it appears in the real world.
NAVER Cloud also unveiled research that further improves AI’s ability to understand images and video. Accepted papers include a study designed to improve recognition of object motion in video (Paper: Why Can’t I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition) and SpatialBoost, which applies text-based reasoning to images of physical spaces, making their three-dimensional structure easier to discern (Paper: SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning). Together, these papers demonstrate TEAM NAVER’s competitive edge in foundational technologies that enhance the accuracy of vision AI systems.
Four papers, including InSpace on 360-degree spatial reconstruction and the Seoul World Model, were selected for ECCV Spotlight, a distinction awarded to roughly 1-2% of all submissions.
Backed by sustained investment in R&D, TEAM NAVER has consistently secured double-digit paper acceptances each year across the three leading global computer vision conferences—ECCV, CVPR (Conference on Computer Vision and Pattern Recognition) and ICCV (International Conference on Computer Vision)—demonstrating its world-class capabilities in spatial intelligence.
TEAM NAVER will continue to invest actively in R&D to secure next-generation AI and spatial intelligence technologies and apply them across a wide range of services, making advanced technology tangible in users’ everyday lives. (END)
[Appendix] TEAM NAVER Papers Accepted to ECCV 2026
1. InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360° Image
2. Wid3R: Wide Field-of-View 3D Reconstruction via Camera Model Conditioning
3. Streaming Dense Voxel Representations for 3D Occupancy Prediction
4. Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints
5. BLASt3R: bundle adjustment of any image set with multi-view matching and monocular priors
6. Whareformer: Learning to track what is where in long egocentric videos
7. Syn4D: a multiview synthetic 4D dataset
8. A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
9. Task Alignment: a simple and effective proxy for model merging in computer vision
10. Human mesh modeling for Any Body
11. Multi-HMR 2: multi-person camera-centric human detection, mesh recovery and tracking
12. Vulnerability of privacy-preserving visual localization methods against diffusion-based attacks
13. IDeaL: data-free multi-teacher distillation via improved dead leaves
14. Autoregressive 3D scene generation from multi-view data
15. Seoul World Model: Grounding World Simulation Models in a Real-World Metropolis
16. TetraSDF: Precise Mesh Extraction with Multi-resolution Tetrahedral Grid
17. On the Reliability of Cue Conflict and Beyond
18. Isotropic Embedding Perturbations for Robust Vision Language Encoders
19. Relaxed Rigidity with Ray-based Grouping for Dynamic Gaussian Splatting
20. Video-Oasis: Rethinking Evaluation of Video Understanding
21. Why Can’t I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
22. Learning from Adversity: Semantic-Aware Mask Refinement through Adversarial Perturbation
23. SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning