<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-20T00:45:50Z</responseDate><request verb="GetRecord" identifier="oai:drum.lib.umd.edu:1903/34629" metadataPrefix="dim">https://api.drum.lib.umd.edu/server/oai/request</request><GetRecord><record><header><identifier>oai:drum.lib.umd.edu:1903/34629</identifier><datestamp>2025-09-15T07:46:19Z</datestamp><setSpec>com_1903_2224</setSpec><setSpec>com_1903_12</setSpec><setSpec>com_1903_2</setSpec><setSpec>col_1903_2756</setSpec><setSpec>col_1903_3</setSpec></header><metadata><dim:dim xmlns:dim="http://www.dspace.org/xmlns/dspace/dim" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://www.dspace.org/xmlns/dspace/dim http://www.dspace.org/schema/dim.xsd">
   <dim:field mdschema="dc" element="contributor" qualifier="advisor" lang="en_US">Shrivastava, Abhinav</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="author" lang="en_US">Huang, Shuaiyi</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="publisher" lang="en_US">Digital Repository at the University of Maryland</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="publisher" lang="en_US">University of Maryland (College Park, Md.)</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="department" lang="en_US">Computer Science</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="accessioned">2025-09-15T05:33:11Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="issued" lang="en_US">2025</dim:field>
   <dim:field mdschema="dc" element="identifier">https://doi.org/10.13016/wb7d-k9fk</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="uri">http://hdl.handle.net/1903/34629</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="abstract" lang="en_US">The rapid growth of visual data and the increasing demand for intelligent robotic systemshave created a pressing need for methods that can establish meaningful correspondences and
relationships across diverse visual modalities and robotic tasks. This dissertation addresses the
fundamental challenge of learning structured alignment, which involves establishing correspondences
between different representations, temporal sequences, and task domains to enable more
effective visual understanding and robot control.

In the first part of this thesis, we advance visual understanding through three key contributionsthat demonstrate the power of structured alignment in perception tasks. We begin by
tackling semantic correspondence, where we propose a teacher-student learning paradigm that
enriches supervision from sparse keypoint annotations, enabling dense correspondence learning
through spatial priors and loss-driven dynamic label selection. We then address video instance
segmentation through two complementary approaches: UVIS, which leverages foundation
models (DINO and CLIP) for unsupervised segmentation without dense annotations, and
PointVIS, which achieves competitive performance using only point-level supervision through
class-agnostic proposal generation and spatio-temporal matching. Finally, we develop Trokens
for few-shot action recognition, introducing semantic-aware point correspondence sampling and
relational motion alignment that captures both intra-trajectory dynamics through Histogram of
Oriented Displacements and inter-trajectory spatial relationships, effectively aligning appearance
features with motion patterns through trajectory-based token alignment.

While the first part focuses on establishing correspondences within visual data, real-worldapplications require bridging the gap between visual understanding and robot control. In the second
part of this thesis, we present two frameworks that demonstrate how structured alignment
can be extended to robotic applications. We introduce ARDuP, a novel method for video-based
policy learning that aligns generated visual plans with language instructions for effective control.
This innovative framework integrates active region (i.e. potential interaction areas) conditioning
with latent diffusion models for video planning and employs latent representations for direct
action decoding during inverse dynamic modeling. By utilizing motion cues in videos for automatic
active region discovery, our method eliminates the need for manual annotations of active
regions. Finally, we present TREND, which addresses robust preference-based reinforcement
learning through a tri-teaching framework that filters noisy preference labels while incorporating
few-shot expert demonstrations, demonstrating effective alignment between human preferences
and robot behaviors even under high noise conditions.</dim:field>
   <dim:field mdschema="dc" element="language" qualifier="iso" lang="en_US">en</dim:field>
   <dim:field mdschema="dc" element="title" lang="en_US">LEARNING STRUCTURED ALIGNMENT: FROM VISUAL UNDERSTANDING TO ROBOT CONTROL</dim:field>
   <dim:field mdschema="dc" element="type" lang="en_US">Dissertation</dim:field>
   <dim:field mdschema="dc" element="subject" qualifier="pqcontrolled" lang="en_US">Computer science</dim:field>
   <dim:field mdschema="dc" element="subject" qualifier="pquncontrolled" lang="en_US">Alignment</dim:field>
   <dim:field mdschema="dc" element="subject" qualifier="pquncontrolled" lang="en_US">Computer Vision</dim:field>
   <dim:field mdschema="dc" element="subject" qualifier="pquncontrolled" lang="en_US">Deep Learning</dim:field>
   <dim:field mdschema="dc" element="subject" qualifier="pquncontrolled" lang="en_US">Generation</dim:field>
   <dim:field mdschema="dc" element="subject" qualifier="pquncontrolled" lang="en_US">Recognition</dim:field>
   <dim:field mdschema="dc" element="subject" qualifier="pquncontrolled" lang="en_US">Robotics</dim:field>
   <dim:field mdschema="others" element="access-status">open.access</dim:field>
</dim:dim>
</metadata></record></GetRecord></OAI-PMH>