Computer Science, 1987-2026
Permanent URI for this collectionhttps://theses-dissertations.princeton.edu/handle/88435/dsp01mp48sc83w
Browse
Browsing Computer Science, 1987-2026 by Author "Abtahi, Parastoo"
- Results Per Page
- Sort Options
A Bit of Clarification: Probabilistic User-in-the-Loop Disambiguation of Multimodal Input in AR
(2026-04-22) Rupertus, Joseph D.; Abtahi, ParastooMultimodal inputs in Augmented Reality (AR) are inherently imprecise, noisy, and underspecified, meaning the same input can map to multiple valid interpretations. This is not a limitation of recognition or inference systems, but rather it is fundamentally an interaction problem. We introduce a proof-of-concept AR system that addresses this through user-in-the-loop disambiguation, using a Dynamic Bayesian Network to fuse gaze, voice, and gesture inputs into probability distributions over target and action spaces. The system enters disambiguation mode when no candidate is sufficiently confident, presenting audio and visual feedback to guide users toward a single clarifying input. We evaluate the system through a two-part user study: a multimodal elicitation study and a disambiguation mode comparison. Our fusion system correctly predicted intent in 53.1% of free-form elicitation trials, with the correct intent appearing in the candidate set in 83.3% of trials. Participants successfully disambiguated after a single input in 82.8% of trials.
Avatar: A Novel Mechanism for Self-Reporting Musically-Evoked Kinesthetic Imagery
(2026-04-25) Liu, Daniel H.; Margulis, Elizabeth Hellmuth; Abtahi, Parastoo; Christianson, KarenMusically-Evoked Kinesthetic Imagery (MEKI) encompasses the internal sensation of dynamic self-movement in response to musical stimuli. While analyzing these sensations offers insights into human cognition and motor response to music, traditional methods for tracking movement are biased towards overt movement, being fundamentally unequipped to measure imagined or physically impossible movements such as floating. Existing self-report methods also lack the dimensional complexity to categorize a broad range of motion features. To address this methodological gap, we are piloting the MEKI avatar: a 3D procedurally-animated humanoid avatar that is web-deployable and designed as a portable self-reporting mechanism that can be used in future MEKI studies. The tool utilizes a system of sliders that isolate specific kinematic parameters including amplitude of movement, speed, and smoothness of motion, while also prioritizing an accessible UI to prevent participant interface fatigue. By translating subjectively experienced MEKI into standardized, quantitative data objects, our avatar offers researchers a high-dimension, flexible new method to visualize and analyze imagined movement.
Implicit Movement-Based Cues as a Trigger for Proactive AR Assistance
(2026-04-16) Huang, Michael; Abtahi, ParastooAugmented reality (AR) assistants leverage spatial understanding and multi-modal input in order to deliver context-aware guidance and instruction. Some systems are proactive, able to fluidly anticipate user needs and predict user goals. However, a key challenge is determining when to intervene, as intervening too early or late can disturb user focus. To address this issue, analysis on the HoloAssist dataset discovered that implicit movement-based cues can serve as a basis for modeling user uncertainty in order to predict ideal moments of intervention. I present a novel AR system that continuously tracks user movement and intervenes at moments of high uncertainty based on these spatial cues. Through visual language model analysis and a decision making algorithm, the system detects movement cues and determines when to proactively give instruction through an AR interface. To evaluate this system, I conducted a within-subjects study (N = 16) involving four physical assembly tasks, measuring user performance and subjective experience. I compared the system against three alternative instructional conditions (Explicit-Asking, Pause, and Step Completion). Contrary to our initial hypothesis, the system was significantly more distracting than the other conditions. Since natural physical exploration heavily overlapped with the movement cues, the system triggered too often with unpredictable interventions, interrupting user focus. However, the system did not significantly degrade task completion time, knowledge retention, perceived workload, or user trust. I discuss the implications of these findings, ultimately proposing guidelines for designing more robust, proactive AR guidance systems.
Item Language-driven Video Editing Agents: Human-AI Creative Collaboration Dynamics Across Interface Modalities
(2024-07-18) Sun, Christine; Abtahi, ParastooDespite the importance of video as a medium for communication, video editing is a skill with a high barrier to entry. To democratize access to content creation, we introduce the Video Editing and Reasoning Agent (VERA), an AI agent that can edit videos based on human instructions. VERA autoregressively conditions on an encoded video representation, viewer context, and command during inference to complete various editing tasks. We show that by equipping the agent with a set of core video editing skills, VERA can reason about combining one or more of them to produce complex edits. To facilitate interaction with the agent, we contribute two interfaces: voice with one-at-a-time feedback and text with aggregated history. We evaluate the VERA system with human participants (N=8) and find statistically significant evidence of the priming effect, in which the initial interface participants are exposed to shapes their subsequent interactions with and perception of the agent. We show that participants who began with the voice-based interface tended to issue commands that were broader in scope and more abstract than those who started with the text-based one. Voice users are also more likely to perceive the agent as a collaborator as opposed to an assistant. Our study's results contribute to the development of human-agent creative collaboration systems that help bring creative vision to life.
Reimagining Home: A New Home Interface Framework for the Apple Vision Pro
(2025-04-10) Kim, Irene; Reinfurt, David; Abtahi, ParastooWith the rise of AR/VR technologies, we are shifting from screen-based computing to spatial computing. In this context, the interface is no longer bounded behind a screen but exists within our space. This thesis questions what it means to design an interface for a space, and reimagines the Apple Vision Pro’s Home View to propose an answer. While the Vision Pro introduces innovative user experiences, its current home screen interface remains rooted in two-dimensional conventions: a window con- sisting of a multi-page grid of flat application icons. Drawing from Apple’s design legacy of simplicity, playfulness, and deference, this thesis introduces a new Home interface framework of two key components: a new visual library of tactile and play- ful 3D application icons and an immersive home space that includes a volumetric App Library and custom interaction model to bring applications to life in the user’s physical space. The resulting interface is one that emphasizes play and personaliza- tion. The interface is evaluated through both a heuristic analysis and scenario-based walkthroughs; through these evaluations we find that the interface’s strength lie in its spatial freedom and user autonomy, and possess opportunities of improvement through diversifying system feedback mechanisms and including a user onboarding. This interface aims to propose a framework for future spatial interfaces and through it, encourage more efforts for research and exploration in spatial UI/UX design.
Item Towards the Ambiguous and Imprecise: Evaluating Feedforward Uncertainty Visualizations for Probabilistic Input Mechanisms
(2024-07-17) Huang, Michelle; Abtahi, ParastooWhile mixed reality interfaces increasingly adopt probabilistic input techniques for more expressive and accessible interactions, these techniques (including eye gaze, voice, and gestures) introduce increased degrees of ambiguity and imprecision. Thus, feedforward mechanisms are key to communicating this uncertainty around the inputs to users in real time. In this paper, I propose an updated design space for feedforward visualizations that communicate ambiguity through various dimensions of visual augmentation, and I evaluate these visualizations in a study with 15 university students to assess their accuracy, confusion, and observations regarding a target selection task informed by the feedforward visualizations. Participants reported that the visualizations that directly modified physical attributes of the targets and the space around the targets were most effective at clearly communicating the presence and source of ambiguity in the target selection to allow participants to adjust their input in real time. I then propose different types of environments and targets in which each visualization would be most suitable, and I conclude by identifying directions for future work with respect to feedforward visualizations in the context of various mixed reality applications.