MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane
arXiv Preprint, 2026
Predicts projected 3D bounding box corners and per-corner depths directly in the image plane without requiring camera intrinsics at inference time.
I am a fourth-year undergraduate student at Yonsei University and an undergraduate research intern with the UCLA Visual Machines Group, where my work is conducted under the supervision of Prof. Achuta Kadambi.
My research interests center on computer vision and generative AI, especially 2D-to-3D lifting, image-plane 3D geometry, and controllable 3D scene generation.
My work focuses on visual 3D understanding from images. I am especially interested in monocular 3D perception, projected image-plane geometry, and generative models that use geometric signals for controllable scene creation.
arXiv Preprint, 2026
Predicts projected 3D bounding box corners and per-corner depths directly in the image plane without requiring camera intrinsics at inference time.