Shanyi Zhang
Papers
1
Total Citations
1
H-Index
1
About
Shanyi Zhang is an emerging researcher in multimodal artificial intelligence, with a focus on weakly supervised learning, affordance grounding, and reasoning-based vision-language models. Their most notable work, "MACR-afford: Weakly supervised multimodal affordance grounding via multi-branch attention enhancement and CoT multi-stage reasoning" (2025), introduces a novel framework that integrates multi-branch attention mechanisms with chain-of-thought (CoT) reasoning to improve the accuracy of affordance detection—the ability to identify how objects can be interacted with—using minimal supervision. This contribution addresses a critical bottleneck in robotics and human-computer interaction, where labeled data is scarce. While early in their career, Zhang’s approach to combining attention enhancement with structured reasoning has already garnered attention, signaling a promising trajectory in advancing interpretable and efficient multimodal systems. Their work stands out for its innovative synthesis of weakly supervised learning and stepwise reasoning, offering a scalable solution for real-world applications like assistive robotics and autonomous navigation. As Zhang continues to build on this foundation, their research is poised to influence the next generation of context-aware AI.
Research Focus
Key Achievements
Top Papers
- 1