Computer Science, 1987-2026
Permanent URI for this collectionhttps://theses-dissertations.princeton.edu/handle/88435/dsp01mp48sc83w
Browse
Browsing Computer Science, 1987-2026 by Title
- Results Per Page
- Sort Options
Item 2.5D Animation in Blender: Combining 2D Tweening, 3D Context, and Physics Simulation Toward 2.5D Simulated Keyframe Animation
(2021-08-17) E, Ilene; Finkelstein, Adam3D animation typically requires specialized skills and training, and tends to limit creative expression in favor of physical feasibility. On the other hand, while drawing2D animation frames is thought to be more accessible and expressive, it is tedious and tricky to inject sophisticated and believable physics into drawings. Another duality within animation exists between forward (simulation-based) animation and pose-to-pose (keyframed) animation. While simulations provide physical believability to the generated animations, keyframes give animators a level of timing control that is missing from the former approach. This project seeks to bridge the gap between these different approaches to animation: leveraging the expressiveness of 2D animation, the robustness of 3D environment and camera movement, the physical feasibility of 3D simulation, and the timing control of keyframing. To this end, I present a 2.5D animation interface that takes 2D drawn keyframes and 3D context (object, environment and camera movement) to generate simulated animations that adhere to the user-drawn keyframes.
Item 3D Model Reconstruction from a "Sea of Images"
(2004) Battaglia, Frank; Funkhouser, ThomasItem 3D Models of Buildings
(1993) McAllister, Jonathan Tillett; Hanrahan, PatrickItem 3d Object Reconstruction of Unseen and Unlabeled Point Clouds
(2022-08-09) Chou, Gene; Heide, Felix; Deng, JiaIn this paper we attempt to reconstruct 3D objects from unseen and unlabeled point clouds. Specifically, we explore generalization capabilities of neural Signed Distance Functions (SDF). 3D object reconstruction is becoming increasingly important for tasks such as self-driving and robotics manipulation, and the ability for a model to reconstruct objects from in-the-wild point clouds with unseen categories is crucial. Previous works using SDFs have achieved impressive results, but they either cannot operate on unlabeled data, or cannot generalize, making their applications limited. We propose a novel semi-supervised setting in which we train on non-overlapping labeled and unlabeled classes. We develop a 2- stage meta-learning approach as well as a self-supervised method to achieve both generalizability and scalability. We evaluate on synthetic and real-world datasets and show that our method outperforms SDF baselines and generalizes to unseen classes with favorable results.
Item 3D Reconstruction with Neural Networks.
(2018-08-14) Melesse, Michael; Rusinkiewicz, SzymonNeural Networks have been successful in tackling problems in computer vision especially over the last few years. Starting with image recognition neural networks, have gone on to produce comparable if not better results in other areas of computer vision such as Image Segmentation and Object Localization. Here we will see an approach in which neural networks can be adapted to deal with another area of computer vision, 3D Reconstruction.
Item 3D Shape Manipulation Using Deep Generative Adversarial Networks
(2017-5-5) Liu, Jerry; Funkhouser, Thomas A.Deep generative models such as Generative Adversarial Networks (GANs) have demonstratedthe ability to effectively learn a manifold over the training data, including 3D data, such thatgenerated objects on this manifold appear crisp and realistic. In the meantime, the process ofcreating detailed, realistic 3D objects by hand is tremendously difficult: it would be incrediblytedious for an unskilled user to create and edit any 3D shape in a realistic fashion. In this work,we propose a voxel-based, comprehensive 3D shape manipulation framework. The frameworkallows users to repeatedly “snap” an imperfect input to a detailed object on the manifold of aGAN, allowing them to create and edit an object with ease. Our framework extends thedefault GAN model by incorporating a projection network and a learned feature space tolearn the snapping operation. We build a shape manipulation application to demonstrate ourresults. Since our main goal is to apply deep learning to a content creation application, wealso assess the general financial impact of 3D deep learning by evaluating whether investorsbehave rationally with respect to deep learning advancements.
Item 3D Surfaces in the Wild
(2019-09-04) Fan, David; Deng, JiaThe recovery of 3D structure from a single 2D image remains an open problem in computer vision. Neural networks do reasonably well at predicting the 3D structure of limited scenes - mostly of indoor scenes and road scenes. But, they are unable to generalize well to unseen training images. We hypothesize that this is in large part due to the lack of diverse and large scale training data for 3D inference. Recent work has attempted to crowdsource 3D annotations of images in the "wild", but due to the large amount of labor involved, fails to produce datasets that are large and expressive enough to improve state-of-art in 3D inference.
Our contribution is three-fold. First, we present a methodology for efficiently obtaining dense 3D annotations of everyday images scraped from the Internet, or images in the wild. Applying this method to Amazon Mechanical Turk workers, we crowdsourced a novel 3D vision dataset of large scale and diversity, which we call "3SIW". We provide full surface normal, depth, fold boundary, and occlusion boundary annotations for 20,000 images from the wild. Our methodology can be used to create other datasets of larger scale and diversity. Secondly, we provide benchmarks on 3SIW for four tasks: surface normal estimation, occlusion detection, fold detection, and semantic segmentation of planar surfaces. Lastly, we demonstrate that training on larger and more diverse data advances the state-of-art in 3D visual systems.
Item 3D Synaptic Cleft Detection of EM Images with Multi-Scale Deep Learning Architectures
(2017-6-3) Lam, Nathaniel; Seung, H. SebastianWith recent advances in electron microscopic (EM) imaging, researchers now have unprecedented access to high resolution image data of neural circuits. Automatic reconstruction of these circuits has shown be vital to the field of connectomics, as manual annotation of EM images is far too time consuming. Current state of the art methods use deep convolutional neural networks (DCNNs) for automatic 3D neuron segmentation reconstruction. In this paper, we leverage current works for the task of synapse detection in adult fly brain, building off of the multi-scale 3D UNet architecture used in [9] for neuron segmentation of EM data. In addition, we demonstrate the effects of gradient masking and multi-task learning with the belief that their application helps remedy challenges specific to the synapse detection problem. We show that the use of multi-task learning, with neuron segmentation as the auxiliary task, can achieve better precision and recall for synapse detection than independent task learning. We argue that reason for improvements result from a shared representation in lower dimensional structure.
A Benchmark for Visual SLAM Based in Infinigen
(2025) Li, Dylan C.; Deng, JiaThe use of synthetic datasets for computer vision is a key factor in the improvement of many methods. Infinigen is a procedural generator of synthetic 3D scenes of the natural world with the goal of creating datasets for computer vision research. One such area of research is Visual SLAM models, which seek to map an unknown environment while simultaneously tracking an agent’s pose. I propose a Visual SLAM benchmark based on Infinigen as well as an open-source method to generate more data for SLAM algorithms using Infinigen. The Infinigen SLAM benchmark contains extremely challenging camera motion within various indoor environments. State-of-the-art Visual SLAM models perform well on the proposed benchmark, however they perform worse than on comparable SLAM benchmarks. This suggests that Infinigen is capable of producing useful data for future SLAM research.
A Bioinformatics Approach to Information-Driven Folding and Docking of Antibody-Antigen Complexes
(2025) Burbank-Embry, Sarah H.; Dieng, Adji BoussoThis thesis presents a user friendly approach to information driven antibody-antigen folding and docking.
A Bit of Clarification: Probabilistic User-in-the-Loop Disambiguation of Multimodal Input in AR
(2026-04-22) Rupertus, Joseph D.; Abtahi, ParastooMultimodal inputs in Augmented Reality (AR) are inherently imprecise, noisy, and underspecified, meaning the same input can map to multiple valid interpretations. This is not a limitation of recognition or inference systems, but rather it is fundamentally an interaction problem. We introduce a proof-of-concept AR system that addresses this through user-in-the-loop disambiguation, using a Dynamic Bayesian Network to fuse gaze, voice, and gesture inputs into probability distributions over target and action spaces. The system enters disambiguation mode when no candidate is sufficiently confident, presenting audio and visual feedback to guide users toward a single clarifying input. We evaluate the system through a two-part user study: a multimodal elicitation study and a disambiguation mode comparison. Our fusion system correctly predicted intent in 53.1% of free-form elicitation trials, with the correct intent appearing in the candidate set in 83.3% of trials. Participants successfully disambiguated after a single input in 82.8% of trials.
A Comparative Study of Syntax and Word Usage Between Standard French and Cameroonian French Using Natural Language Processing
(2025-04-10) Hines, Julia R.; Fellbaum, Christiane DorotheaThis study uses natural language processing (NLP) techniques to analyze the syntactic and lexical differences between Standard French and Cameroonian French, as well as examine how the dialect evolves when used by the Cameroonian diaspora in France. The central methodology involves training and evaluating two distinct NLP models: one fine-tuned on a corpus of Standard French, and the other on Cameroonian French. The LSTM model, on the other hand, outperformed the Logistic Regression model in all key metrics, including accuracy, precision, recall, and F1-score. The results of this study illustrate the limitations of traditional NLP methods, such as logistic regression, when applied to dialects with syntactical and linguistic differences, and they highlight the potential of deep learning approaches to better handle these variations. The findings point to the importance of fostering linguistic diversity within computational models.
A Comparison of Model Predictive Control and Reinforcement Learning Methods for Building Energy Storage Management
(2025-04-10) Toh, Yi Jin; Eysenbach, BenjaminThe residential building sector is a major contributor to energy consumption and greenhouse gas emissions, making electrification and intelligent energy management essential for decarbonization. However, increased electricity demand can strain the power grid, leading to higher costs and emissions. Demand-side flexibility, enabled by on-site power generation, energy storage, and optimized control algorithms, can mitigate this problem by shifting electricity consumption to times when electricity is cheaper and cleaner.
This study evaluates three methods for centralized building energy storage management using CityLearn, an open-source environment for simulating and benchmarking building energy control. The evaluation compares Model Predictive Control (MPC) with two Reinforcement Learning (RL) methods: Soft Actor-Critic (SAC) and Proximal Policy Optimization (PPO). The methods are assessed across three dimensions: (1) energy performance, including cost, carbon emissions, electricity consumption, and stability of electricity use over time; (2) computational efficiency, including training time, memory usage, and inference speed; and (3) scalability, measured across different district sizes of two, four, and eight buildings.
Overall, SAC achieved the strongest performance on cost and energy metrics, performing slightly better than PPO in those areas. PPO, however, produced smoother control behavior with more stable electricity use over time while requiring significantly less memory than SAC and less computation than MPC. Both RL methods outperformed MPC across most metrics, with MPC particularly struggling to scale. Nonetheless, MPC remained more interpretable and required no training data, though it involved substantial engineering effort to develop an accurate system model.
These findings highlight trade-offs between performance, stability, and deployability. PPO emerged as the most balanced controller, offering strong performance with scalability and computational efficiency, making it well-suited for real-world use.
A Computational Model of Intertemporal Choice: Exploring the Impact of Sleep Deprivation on the Discounting of Future Rewards
(2025) Botton, Estelle; Niv, YaelSleep has a profound impact on numerous cognitive functions, including decision-making and self-control. However, the precise mechanisms by which sleep influences these processes are not completely understood. This study explored how sleep loss impacts the weighting of short-term and long-term rewards, and thereby influences the decision-making process. Participants self-reported their sleep on the previous night and participated in a decision-making experiment involving 25 choices between pairs of food items, where the taste value of each food item represented short-term reward and the health value of the food represented long-term reward. I hypothesized that participants who slept less would exhibit less self-control and thus would place a lesser weight on health relative to taste; further, I hypothesized that self-control would deplete as trials progress.
I developed four nested computational models to characterize the decision-making process: a baseline model that assumes equal weights for taste and health, a model that fits a health weighting parameter for each participant, and two models that incorporate a linear or exponential decay parameter to simulate potential self-control depletion across trials. The model that fit a βhealth weight for each participant provided the best fit for the data, suggesting that participants vary in how they weigh taste versus health and that this weighting was relatively stable across trials. This study did not find a significant correlation between βhealth and sleep hours under any of the models. While the results did not align with my hypotheses, this may be due to a small sample size, limited variability in sleep duration, computational constraints, and other factors. Further research is needed to better explore this relationship, potentially with more extreme manipulations of sleep or the consideration of additional factors that may influence intertemporal choice.
A Computational Study of Persuasion in Dialogue: Linguistic Features and Conversational Context
(2026-04-16) Ali, Laiba; Bhat, Suma PallathadkaThis work analyzed the relationship between persuasive language and its observable effect, agreement, in structured dyadic dialogue. In particular, we studied the presence and distribution of linguistic features and Cialdini-inspired persuasion techniques regarding conversational behavior. We leveraged large language models (LLMs) to annotate for persuasion principles across multiple dialogue datasets, while also examining the reliability and consistency of LLM-based evaluation. We analyzed these datasets using a combination of Ordinary Least Squares (OLS) linear regression modeling, Multivariate Analyses of Variance (MANOVAs), K-Means clustering, and descriptive comparisons to evaluate whether these features were indicative of conversational outcomes and contextual conditions, specifically in the form of pre-conversational prompting. Our findings showed that while these features provided limited predictive power for agreement outcomes, the discourse markers, structural properties, and persuasion principles varied across different prompting conditions: persuasion, compromise, and general dyadic dialogue. These results suggest that conversational intent played a significant role in shaping feature usage, while also emphasizing the importance of context in computational analyses of dialogue.
A Computer Vision Approach to Analyzing Player Movement
(2025) Aguirre, Maria F.; Heide, FelixThis thesis provides a new resource for squash performance analysis by developing a computer vision system that integrates advanced object detection and tracking techniques. By stringing together YOLOv8 for precise player detection and StrongSORT for multi-object tracking, the system accurately processes game footage to collect player data. A tool developed in this project is a user-assisted manual court mapping interface corrects perspective distortions, providing the resource to generate movement based analytics that reflect on-court dynamics, such as control of the critical ’T’ position. The adaptability of the technologies used create the opportunity for the expansion of this project. Further development of this project offers valuable insight for coaching and performance improvement, and further refinements are expected to enhance the accuracy of detection and tracking even further.
A Content-Aware Time Compression Algorithm for Audio - CATCA
(2025) Hoffman, Ryan F.; Finkelstein, AdamModern audio time-compression algorithms generally follow a uniform approach to speedup. Given a particular playback rate, these algorithms decrease the number of audio samples played evenly throughout the entire clip and use a variety of techniques to control the pitch so that it remains constant. This is generally e!ective until higher speeds, past which the quality of the audio degrades to a point of lacking comprehensibility to the listener. However, by designing an algorithm that analyzes the frequencies in each audio sample and removes them strategically according to their perceived importance, it is theoretically possible to preserve the intelligibility of an audio file better even at higher playback rates. This unlocks potentially higher speeds for listener comprehension and improves the listening experience at standard playback rates. This algorithm, called CATCA (Content-Aware Time Compression for Audio), is built on a content-aware approach, which assigns energies to audio samples and removes them in priority of lowest energy. While this new time-compression algorithm did not achieve intelligibility improvements over the state-of-the-art method PSOLA, it still performed better than other algorithm variations, demonstrating the utility of content-awareness as an audio time-compression approach approach given future improvements.
A Corpus-Based Approach to English Adversative Coordination
(2025) Weizel, Oliver L.; Fellbaum, Christiane DorotheaWhat is the difference between and and but? In many sentences, they can be freely interchanged—consider that both ”the weather is sunny and cold” and ”the weather is sunny but cold” are true if and only if the weather is sunny and the weather is cold. What then causes speakers to chose and over but and vice versa? To that end, I gather data from the Corpus of Contemporary American English and investigate properties of the distributions of the two conjunctions. I find that and is more unmarked and neutral, while but is more likely to appear when greater contrasts exist between the two conjuncts themselves, or more broadly in more salient contexts. Along the way, novel analyses for the underlying syntactic structure of certain uses of but are proposed.
A Generative AI-Based End-to-End Pipeline for Game-Style Visualization of Journey to the West
(2025) Liu, Annie; Kernighan, Brian W.This project proposes an automated end-to-end pipeline that transforms Journey to the West into a game-style visualization. Leveraging generative AI models for both text and image, the pipeline extracts structured data from the original narrative and converts it into consistent, stylized visual assets. These components are then integrated within an interactive interface to visualize the story. By fully automating the process, the project lowers the barrier to entry for new forms of engagement with classical literature, while also exploring the limits of prompt engineering and highlighting both the creative potential and structural challenges of using AI to reinterpret complex literary works.
A Lucid Solution: Online Machine Learning for DDoS Attack Detection in Programmable Switches
(2026-04-16) Kim, Grace; Walker, David P.Distributed Denial-of-Service (DDoS) attacks are cyberattacks in which a malicious actor floods a service, server, or network with illegitimate internet traffic. The malicious actor uses multiple systems, such as a network of devices with unique IP addresses, to carry out this distributed attack. DDoS attacks prevent legitimate users from being served, and compromise the performance and reliability of modern computer networks such as cloud datacenters. Machine Learning (ML) is increasingly being used in DDoS attack detection, which aims to distinguish malicious DDoS traffic from benign traffic.
We identify three key challenges faced by in-network systems that use ML for DDoS attack detection. First, it is critical that the system has both high performance and low latency. Second, the system must continuously monitor for concept drift: the event in which the ML model in the system becomes ineffective at detecting attacks due to massive changes in the traffic distribution. The ML model must be updated if concept drift occurs. Third, these in-network systems often use languages that are widely known to be difficult and verbose, such as P4, which leads to delays in development.
We propose a system design which would use Lucid, a Domain Specific Language (DSL) developed in Princeton, to implement a decision tree on programmable switches. This system would utilize 2-Way D-Left hashing for stateful feature tracking, an SRAM array to encode the decision tree's rules, and concept drift detection directly on the switch using an algorithm adapted from Exponentially Weighted Moving Average (EWMA). Our Proof-of-Concept system utilizes a pruned decision tree of 92% test accuracy and 13 nodes. We demonstrate that our Lucid implementation is compatible with standard programmable switches. Our Lucid implementation requires 20 ~ 30 times fewer lines of code than the equivalent implementation in P4.