NewsVisual positioning

AI-powered visual positioning can find you from just one photo.

MultiSet matched an ordinary photo against a 3D scan of Austin. The visible scene led it back to the camera’s exact position.

ShareXFacebookLinkedIn
A 3D scan of Austin with colored lines connecting photo features to their matched positions in the city
Photo features are matched to the same points inside a 3D scan of Austin.Image: Bilawal Sidhu and MultiSet via X

Bilawal Sidhu took several photos around Austin and tested them with MultiSet’s private visual positioning system (opens in a new tab). One came from a high-rise window. The system did more than recognize the skyline. It found the building, the height, and the exact balcony where the camera had been standing.

The buildings, roads, and skyline became clues that led back to the camera. The view itself became location data.

How it found the balcony

To understand how, it helps to think of visual positioning as GPS for cameras. First, a place is scanned and turned into a 3D map. A new photo can then be compared with that map to find where the camera is and which way it is pointing.

Sidhu’s demo makes that process visible. Colored lines connect details in each Austin photo to the same points inside the city scan. Once enough details match, the photos and video frames snap back into the places where they were captured.

What made it possible

That precision did not come from the photo alone. The test worked because that part of Austin had already been scanned and Sidhu could access the matching map. This is not a public tool that can identify every home from any photo.

That boundary matters. MultiSet says its maps are private to each customer account (opens in a new tab), so other developers cannot search them as one shared public world map.

For now, that limits how widely a photo can be matched. Even so, detailed maps are becoming easier to create. MultiSet can use phone LiDAR, 360-degree video, point clouds, Matterport scans, textured models, and Gaussian splats. As more places are mapped, more ordinary scenes could be recognized this way.

Why the view matters

The experiment changes how we should think about photo privacy. Removing GPS metadata may not be enough when the buildings, roads, windows, and skyline form their own visual fingerprint.

A skyline photo can seem anonymous on its own. When it is compared with a matching map, the same view could point to a home, hotel, or office, then narrow the location to a building, floor, and window.

The same matching process also has useful applications. It can keep digital directions attached to the right door or help a robot locate itself indoors. That is the tension at the heart of visual positioning. The match can make spatial apps and robots more capable, while the location it reveals remains sensitive.

Today, not every skyline photo exposes an address. Once a place has been mapped, the pixels in the view may be enough.

More from Capture & Create.

Browse all stories
GOLF+ promotional image of a Quest wearer putting a real golf ball in a living room

GOLF+ · 2 min read

GOLF+ Now Lets You Putt With a Real Club and Ball on Quest

Immersive Putting brings virtual greens into your room, using the speed and direction of a real ball to simulate where your shot would go.

Three Hyundai development vehicles on display at the Group’s Autonomous Driving Media Day

Autonomous Driving · 2 min read

Hyundai Recreates Real Roads in 3D to Test Its Driving AI

Hyundai uses Gaussian splatting to replay difficult driving scenarios in simulation, alongside plans to bring NVIDIA-powered driver assistance to production cars in 2028.

XGRIDS Lixel L3 handheld scanner with its cameras and LiDAR sensor against a black background

3D Scanning · 2 min read

XGRIDS Unveils Lixel L3 Scanner With 5 mm Accuracy and 4K HDR

The L3 upgrades XGRIDS’ handheld scanning hardware with larger camera sensors, built-in positioning, and a screen for use in the field.

Preparing the Spatial Insider studio