Paper Reading AI Learner

REACT++: Efficient Cross-Attention for Real-Time Scene Graph Generation

2026-03-06 15:36:00
Ma\"elic Neau, Zoe Falomir

Abstract

Scene Graph Generation (SGG) is a task that encodes visual relationships between objects in images as graph structures. SGG shows significant promise as a foundational component for downstream tasks, such as reasoning for embodied agents. To enable real-time applications, SGG must address the trade-off between performance and inference speed. However, current methods tend to focus on one of the following: (1) improving relation prediction accuracy, (2) enhancing object detection accuracy, or (3) reducing latency, without aiming to balance all three objectives simultaneously. To address this limitation, we build on the powerful Real-time Efficiency and Accuracy Compromise for Tradeoffs in Scene Graph Generation (REACT) architecture and propose REACT++, a new state-of-the-art model for real-time SGG. By leveraging efficient feature extraction and subject-to-object cross-attention within the prototype space, REACT++ balances latency and representational power. REACT++ achieves the highest inference speed among existing SGG models, improving relation prediction accuracy without sacrificing object detection performance. Compared to the previous REACT version, REACT++ is 20% faster with a gain of 10% in relation prediction accuracy on average. The code is available at this https URL.

Abstract (translated)

场景图生成(SGG)是一项任务,旨在将图像中对象之间的视觉关系编码为图结构。SGG作为下游任务的基础组件表现出巨大潜力,例如用于具身代理的推理。为了实现实时应用,SGG必须解决性能和推断速度之间的权衡问题。然而,目前的方法往往集中于以下三个目标之一:(1) 提高关系预测准确性;(2) 增强对象检测准确性;或 (3) 减少延迟,并没有同时兼顾这三个目标。 为了解决这一局限性,我们基于强大的“实时效率与精度权衡的场景图生成架构”(REACT)构建了 REACT++,这是一种新的、面向实时SGG的状态-of-the-art模型。通过利用高效的特征提取和原型空间内的主客体交叉注意机制,REACT++ 平衡了延迟和表示能力之间的关系。REACT++ 达到了现有 SGG 模型中最快的推断速度,在不牺牲对象检测性能的情况下提高了关系预测准确性。 与之前的 REACT 版本相比,REACT++ 在关系预测准确性的平均值上提升了10%,并且速度快20%。代码可在该网址获得:[此链接指向的是原文中的URL位置,请根据实际情况提供或省略]。

URL

https://arxiv.org/abs/2603.06386

PDF

https://arxiv.org/pdf/2603.06386.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot