- Updated: August 22, 2026
- 1 min read
HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation
HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation
We are excited to present HSTGFormer, a cutting‑edge graph‑enhanced Transformer framework that advances monocular 3D human pose estimation. By unifying spatial‑temporal reasoning through a Hyper Spatial‑Temporal Graph (HSTG), HSTGFormer captures local structural motion information while preserving global context.
The core innovations include:
- Hyper Spatial‑Temporal Graph (HSTG): Extends per‑frame skeleton graphs into temporal neighborhoods, enabling coupled spatial‑temporal aggregation around each joint‑time node.
- Adaptive Dual‑Scale Temporal Graph (ADSTG): Learns joint‑specific short‑ and long‑range temporal dependencies for robust motion modeling.
- Node‑wise Fusion Module: Dynamically merges HSTG and ADSTG representations for each joint‑time node.
Extensive experiments on Human3.6M and MPI‑INF‑3DHP demonstrate that HSTGFormer achieves state‑of‑the‑art accuracy with high computational efficiency, making it ideal for real‑time applications.
Read the full arXiv paper for technical details: https://arxiv.org/abs/2608.12187.
Explore related research and implementation resources on our site:
- Graph Transformers
- HSTGFormer Architecture Overview
- Contact Us for collaborations.
Stay tuned for upcoming tutorials and open‑source code releases.
Andrii Bidochko
CTO UBOS
Andrii Bidochko is an AI entrepreneur and researcher focused on AI agents, reinforcement learning, and autonomous systems. He writes about the technologies shaping the future of machine intelligence, from frontier models and agent architectures to real-world AI applications.