Seekr



Seekr is a desktop research prototype that combines Gaussian Splatting environments, embodied AI agents, and natural-language interaction to explore how humans and AI perceive and represent space differently. I designed and implemented the frontend experience, allowing users to initiate agent exploration through chat, track its trajectory and camera orientation, and review the perception frames and object data returned by the AI system. I also contributed to the project’s visual and spatial development through character design and modeling, world generation, and mesh cleanup. Seekr became the starting point for my ongoing research into personalized world models and spatial cognition.

Qwen, Gaussian Splatting, AI World Models, LLM, Tripo
Collaborative Research










Seekr operates through a two-phase conversational interaction. First, the user enters a command such as “Explore the surroundings,” prompting the AI explorer agent to control a character and autonomously navigate the 3D environment.

After the exploration, the user can ask, “Describe what you see.” Seekr then captures an image from the character’s current viewpoint and sends it to the Perception Brain, where the Qwen vision-language model analyzes the scene. The system identifies visible objects, interprets the spatial context, and returns a description of the environment through the chat interface.






Seekr supports Gaussian Splatting environments generated with Marble and Habitat-based environments. 

* Habitat is a benchmark developed by Meta, a simulation platform for evaluating embodied AI agents on navigation, perception, and interaction tasks in realistic 3D environments.







[Left] Map and Chat tabs for trajectory tracking, camera orientation, and natural-language interaction with the agent

[Right] Object tab displaying objects identified by the Perception Brain
Perception view captured by the agent
The Seekr interface brings navigation, communication, and visual perception into a unified workspace. The Map view tracks the agent’s trajectory and camera orientation, while the Chat tab enables users to guide the agent through natural-language commands. When prompted to describe its surroundings, the agent captures its current view and sends it to the Perception Brain, which returns a spatial description and a list of identified objects stored in the Objects tab.



Click to scroll to top
ⓒ 2026. MinyoungJoo Phenomenological Design ResearcherLast Updated September 2026