We describe a monocular, pedestrian-specific single-stage architecture in which one transformer encoder serves both detection and short-horizon trajectory prediction, and we position it explicitly against the existing joint perception-and-prediction literature rather than around it.
We introduce a Spatial-Temporal Decoder that composes pedestrian self-attention, scene cross-attention, and temporal self-attention inside each decoder layer, so that crowd context, visual grounding, and motion history are mixed before either head reads the query.
We show that a four-term objective combining focal, smooth-$L_1$, ADE, and FDE losses converges without the loss balancing problems that often affect multi-task training, and we report the resulting behaviour honestly, including the synthetic nature of the benchmark and the metrics it does not support.
