STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction

CoRL 2026

Nathan Tsoi*✉, Michael J. Munje*, Tejas Oberoi, Rishab Maheshwari, Pengen Zheng, Tanush Chauhan, Peter Stone, Joydeep Biswas
Department of Computer Science, The University of Texas at Austin
* Equal contribution
✉ Correspondence: nathan.tsoi@utexas.edu

Overview

Overview of the three STARS stages: graph construction, unsupervised pretraining, and linear probing
The three stages of STARS: (A) construct spatiotemporal graphs from multi-modal Human-Robot Interaction datasets, (B) use an unsupervised message-passing graph autoencoder to compress physical dynamics into latent representations, and (C) adapt the frozen latents to downstream tasks with lightweight task-specific prediction heads.

Abstract

Fielding socially competent robots requires joint reasoning over spatial information and the social information it can convey, such as group motion and personal space. While current models excel at either spatial reasoning (e.g., modeling physical dynamics) or social reasoning (e.g., parsing text-based interactions), few natively unify both. Moreover, data scarcity in Human-Robot Interaction (HRI) makes training large models from scratch challenging. To bridge this gap, we introduce STARS: SpatioTemporal Autoencoded Representations for Social-interaction, a method for learning event-level relational representations that summarize short interaction scenarios from social navigation datasets. Rather than training task-specific models, STARS models fixed-duration interaction scenarios, i.e., short windows of multi-agent interaction, as spatiotemporal graphs. A self-supervised message-passing graph neural network autoencoder then compresses these physical dynamics into a compact latent space for each node. By explicitly modeling agents and their interactions as graph structure, STARS injects a strong relational inductive bias to naturally capture the underlying social situation. To assess the utility of these learned representations, we evaluate STARS' representations on the SEAN Together dataset, which provides VR-based navigation data annotated with subjective human perceptions, and the SocialNav-SUB benchmark for social scene understanding. When trained on STARS representations, linear probes achieve highly competitive performance on perception prediction and pedestrian action classification tasks. Specifically, STARS exhibits strong data efficiency, with its largest gains over unstructured baselines on pedestrian action classification, where STARS-He achieves a macro F1-Score of 0.362 with 1% of labeled data and 0.503 with 100%. Furthermore, qualitative analysis of the learned latent space reveals that STARS representations can capture high-level social semantics, demonstrating cohesive clustering of distinct human action intents without optimizing the encoder with a supervised action classification objective.

What's in STARS

A

Interaction scenarios as graphs

Each fixed-duration window of multi-agent motion becomes a complete, edge-feature-dense graph: agents are nodes and every ordered pair carries a temporal sequence of relative SE(2) poses and their deltas (8 channels). Datasets with different sensors, rates, and annotations map into one schema.

B

Self-supervised graph autoencoder

A message-passing GNN-VAE encodes each node to a 128-D Gaussian latent and reconstructs the temporal edge features from latent pairs. No labels are used during pretraining, so the encoder can be larger than a supervised model of the same architecture without overfitting the scarce annotations.

C

Frozen latents, linear probes

The encoder is frozen and a linear probe reads out task-relevant latents — robot-plus-mean-human for graph-level SEAN-T perception, robot-plus-target-human for node-level SNS action classification. Downstream scores therefore reflect the representation, not task-specific representation learning.

We instantiate STARS in two variants: STARS-Ho (homogeneous graph, one node and edge type) and STARS-He (heterogeneous graph, typed robot/bystander nodes and typed relations).

Datasets & Learned Latent Space

Left: example frames from the SEAN-T and SocialNav-SUB datasets. Right: PCA and t-SNE projections of STARS human-node latents colored by avoid and follow actions.
(A) Dataset overview. We use two human-robot interaction datasets: SEAN-T for subjective human ratings of robot navigation, and SNS (SocialNav-SUB) for relational pedestrian action classification (e.g., avoid, follow). (B) Latent space analysis. 2D projections of the 128-dimensional SNS human-node embeddings from the STARS-trained homogeneous GNN-VAE encoder, using linear PCA and non-linear t-SNE, color-coded by action label. The encoder was never optimized with a supervised action classification objective, yet the latent space self-organizes by action intent.

Dataset and graph construction

Dataset # Scenarios Sample freq. Window size T Window duration Max nodes N Max edges
SEAN-T 2969 5 Hz 40 8 s 16 240
SocialNav-SUB 3052 8 Hz 20 2.5 s 31 930

The temporal window is fixed within each dataset; the maximum number of edges follows from the maximum number of nodes under the fully connected graph construction.

Downstream Performance and Data Efficiency

We pretrain STARS on the combined training splits with the self-supervised reconstruction objective, freeze the encoder, and train a linear probe on progressively larger fractions of the labels. The clearest gap appears on SNS pedestrian action classification: unstructured baselines plateau near a macro F1-Score of 0.3 no matter how much labeled data they receive, while STARS-He reaches 0.362 with just 1% of the labels (60 samples) and climbs to 0.503 at 100%.

SNS pedestrian action classification — macro F1 vs. labeled-data fraction

0.30 0.35 0.40 0.45 0.50 1% 5% 10% 25% 50% 100% Fraction of labeled training data STARS-He STARS-Ho Autoencoder MLP
STARS-He (ours) STARS-Ho (ours) Autoencoder MLP
Macro F1-Score on the held-out SNS action classification test set, averaged over 10 seeds. Full numbers, including standard deviations, are in the table below.

Full results

Task Method 1% 5% 10% 25% 50% 100%
SEAN-T
Competence
MLP 0.450 ± 0.220.546 ± 0.190.664 ± 0.150.761 ± 0.030.778 ± 0.020.797 ± 0.02
Autoencoder 0.663 ± 0.080.678 ± 0.050.706 ± 0.050.734 ± 0.020.747 ± 0.020.742 ± 0.02
Random Forest 0.557 ± 0.200.641 ± 0.120.680 ± 0.100.704 ± 0.090.715 ± 0.090.718 ± 0.08
STARS-He (ours) 0.635 ± 0.100.673 ± 0.070.726 ± 0.040.762 ± 0.030.781 ± 0.030.794 ± 0.02
STARS-Ho (ours) 0.612 ± 0.140.702 ± 0.080.738 ± 0.050.765 ± 0.040.787 ± 0.030.799 ± 0.02
SEAN-T
Surprise
MLP 0.217 ± 0.220.250 ± 0.140.423 ± 0.270.450 ± 0.210.630 ± 0.230.720 ± 0.09
Autoencoder 0.548 ± 0.100.597 ± 0.120.660 ± 0.070.674 ± 0.090.712 ± 0.100.739 ± 0.06
Random Forest 0.291 ± 0.240.341 ± 0.190.373 ± 0.200.444 ± 0.160.470 ± 0.150.496 ± 0.15
STARS-He (ours) 0.455 ± 0.170.481 ± 0.170.526 ± 0.150.603 ± 0.120.655 ± 0.080.709 ± 0.04
STARS-Ho (ours) 0.440 ± 0.230.541 ± 0.210.548 ± 0.200.610 ± 0.160.713 ± 0.060.741 ± 0.03
SEAN-T
Intention
MLP 0.479 ± 0.200.533 ± 0.130.606 ± 0.060.659 ± 0.030.693 ± 0.030.703 ± 0.01
Autoencoder 0.569 ± 0.040.622 ± 0.030.659 ± 0.030.665 ± 0.050.661 ± 0.030.664 ± 0.02
Random Forest 0.592 ± 0.180.665 ± 0.100.685 ± 0.090.692 ± 0.090.700 ± 0.080.698 ± 0.09
STARS-He (ours) 0.576 ± 0.080.619 ± 0.050.671 ± 0.040.694 ± 0.030.705 ± 0.030.721 ± 0.03
STARS-Ho (ours) 0.603 ± 0.100.671 ± 0.070.701 ± 0.060.715 ± 0.040.723 ± 0.030.734 ± 0.02
SNS
Macro F1
MLP 0.294 ± 0.000.297 ± 0.010.300 ± 0.010.306 ± 0.020.295 ± 0.010.293 ± 0.00
Autoencoder 0.308 ± 0.0190.309 ± 0.0180.307 ± 0.0220.310 ± 0.0180.309 ± 0.0180.332 ± 0.008
STARS-He (ours) 0.362 ± 0.0580.447 ± 0.0510.470 ± 0.0440.493 ± 0.0320.501 ± 0.0280.503 ± 0.022
STARS-Ho (ours) 0.344 ± 0.050.431 ± 0.0450.453 ± 0.0350.468 ± 0.030.471 ± 0.0240.475 ± 0.02

F1-Scores (μ ± σ) over 10 random seeds, across progressively scaled fractions of labeled data. STARS-He is the heterogeneous variant and STARS-Ho the homogeneous one; both use a frozen pretrained encoder with a lightweight linear probe. Best score in each column of a block is in bold. Random Forest is omitted for SNS because of the dataset's variable-sized, high-dimensional features.

Further results reported in the paper
  • Self-supervised vs. supervised. The identical heterogeneous architecture trained end-to-end on SNS labels reaches 0.445 macro F1 at 100% labeled data, versus 0.503 for frozen STARS-He plus a linear probe.
  • Cross-dataset transfer. STARS-He pretrained only on SEAN-T reaches 0.497 macro F1 on SNS (vs. 0.503 under joint pretraining); pretrained only on SNS it reaches 0.801 / 0.710 / 0.729 on SEAN-T competence / surprise / intention (vs. 0.794 / 0.709 / 0.721 jointly).
  • Where relational structure matters. STARS-He beats the temporal Autoencoder at every labeled-data fraction on SNS, while on SEAN-T the Autoencoder stays competitive — explicit relational structure helps pedestrian action classification more than perception prediction.

Architecture

STARS heterogeneous architecture: raw inputs, preprocessing, HEAT message-passing encoder, latent space, and decoder
STARS heterogeneous architecture. Raw inputs (robot occupancy grid and goal, learned human embeddings, 8-D temporal edge features) are projected to a 256-D shared hidden space by type-specific preprocessing — MLPs for nodes, a Transformer for edges. The encoder performs HEAT-style message passing over the fully connected heterogeneous graph, with attention and messages parameterized over the concatenated edge-conditioned representation uij. The resulting 256-D node states are projected through shared linear heads to per-node Gaussian posteriors and sampled to 128-D latents zi via the reparameterization trick. The decoder reconstructs edge features (and robot features on SEAN-T) from latent pairs, trained jointly with a KL divergence term against N(0, I). Pretrained encoders are frozen for downstream linear probing.

BibTeX

@article{tsoi2026stars,
  title   = {STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction},
  author  = {Tsoi, Nathan and Munje, Michael J. and Oberoi, Tejas and Maheshwari, Rishab
             and Zheng, Pengen and Chauhan, Tanush and Stone, Peter and Biswas, Joydeep},
  journal = {Proceedings of The 10th Conference on Robot Learning},
  year    = {2026}
}