# Hiiiiiii, I am Ruyi (Maggie) Zhang
I am a PhD student in Computer Science and Technology at the Harbin Institute of Technology (HIT), supervised by Prof. Wangmeng Zuo. Before that, I spent six happy years at the University of Manchester, where I received a BSc in Artificial Intelligence (First Class Honours) and an MSc in Robotics (Distinction).
My research focuses on building multimodal models that understand, generate and generalize in open-world scenarios:
- Multimodal Large Language Models – fine-grained understanding of academic figures, tables and documents
- Semantic Alignment – aligning vision, audio and language across heterogeneous modalities
- Open-Scene Generalization – zero-shot transfer of LMMs to unseen visual modalities (thermal, depth, X-ray, …)
- 3D Scene Generation – controllable text-to-3D scene generation with structured layout priors
Multimodal LLM Semantic Alignment Open-World Generalization 3D Scene Generation Robotics
I am always happy to chat about research and collaboration. Feel free to reach me at 3242701514@qq.com or find me on GitHub.
# News
- 2026 – Two papers accepted to ECCV 2026: VVM-Tuning (generalizing LMMs to versatile visual modalities) and ScrollScape (32K image generation with video diffusion priors).
- 2026 – First-author paper ControlScene accepted to CAAI Transactions.
- 2026 – Paper on task-uncertainty-aware video restoration accepted to IEEE TMM.
- 2025.03 – Started my PhD at the Harbin Institute of Technology.
- 2024.12 – Graduated from the University of Manchester with an MSc in Robotics (Distinction).
# Education
Ph.D. in Computer Science and Technology · Harbin Institute of Technology · 2025.03 -- Present
- Supervisor: Prof. Wangmeng Zuo
- Research: multimodal large language models, semantic alignment, open-scene generalization, 3D scene generation
M.Sc. in Robotics · University of Manchester · 2022.09 -- 2024.12
- Degree: Distinction
- Dissertation: Visual perception for underwater robots with diffusion-model image priors
B.Sc. in Artificial Intelligence · University of Manchester · 2018.09 -- 2022.06
- Degree: First Class Honours
- Honour: Best Undergraduate Dissertation
# Research Experience
Multimodal Intelligence Lab, Harbin Institute of Technology · PhD Researcher · 2025.03 – Present
- Leading research on multimodal understanding of academic literature. Designed the Count-Curriculum framework, which tackles weak gradients on hard samples and raises figure-matching accuracy by 23.3%.
- Proposed ControlScene, a structured-layout-guided 3D scene generation method, and built the LayoutVerse benchmark with 20,000 manually annotated samples.
- Exploring semantic alignment mechanisms of multimodal large models in open scenarios, forming an “understanding – generation – generalization” research line.
Robotics Lab, University of Manchester · MSc Researcher · 2022.09 – 2024.12
- Studied visual perception for underwater robots; developed diffusion-model-based image prior modeling that notably improves detection robustness in low-light and turbid water.
- Designed and optimized a Region-of-Interest Vision (RoIV) model with external attention for real-time object detection under limited compute.
Biological Machine Learning Project, MIT · Research Assistant (Remote) · 2021.06 – 2021.09
- The youngest member of the team, responsible for data preprocessing and feature engineering for deep learning on biological sequence data.
- Helped design a multi-task autoencoder for translating between scRNA and scATAC data; the project ranked first in the group.
Heilongjiang HIT AI Industry Co., Ltd. · Algorithm Engineer Intern · 2022.06 – 2022.09
- Optimized CNN architectures for image recognition via hyper-parameter tuning and model distillation.
- Fine-tuned NLP models for textual semantic understanding in industrial data analysis.
# Selected Projects
- Autonomous Robot Design & Assembly (MSc core project, 2023.09 – 2024.02) – Designed and built an autonomous robot integrating path planning and computer vision for navigation and precise robotic-arm manipulation in complex environments; reduced control latency to improve responsiveness in dynamic scenes.
- Comparative Analysis of Dimensionality Reduction (BSc dissertation) – Systematically evaluated PCA, LDA, t-SNE and UMAP on noisy high-dimensional data and showed UMAP’s advantage in preserving local structure. Awarded Best Undergraduate Dissertation.
- Hex AI Bot (2022.11) – Co-developed an evaluation strategy based on FastVC (virtual connection) search with heuristic rules and a swap-opening strategy, reaching about 75% win rate against other AI agents.
- More course projects are on the home page – chess, maze solver, snake, fuzzy-logic washing machine and more.
# Honours & Awards
- MSc with Distinction; BSc with First Class Honours; Best Undergraduate Dissertation
- Huawei Software Elite Challenge 2022 – National rank 71 (top 3%)
- 2020 Public Welfare Website Design Competition – Best Creativity Award
- PASS Leader, University of Manchester; Outstanding Undergraduate Student Representative
# Skills
–|--
Programming | Python, PyTorch, TensorFlow, C++, Bash, SQL, LaTeX
Multimodal / LLM | Qwen-VL, LLaVA, ImageBind, CLIP, BLIP-2, LoRA / QLoRA
3D / Graphics | 3D Gaussian Splatting, NeRF, Open3D, PyTorch3D
Tools | Linux, Git, Docker, Slurm, Weights & Biases, OpenCV
Languages | Chinese (native), English (fluent, IELTS 7.0, BSc & MSc taught in English)
# 你好呀,我是张茹怡 (Maggie)
我目前在哈尔滨工业大学计算机科学与技术专业攻读博士学位,导师是左旺孟教授。在此之前,我在曼彻斯特大学度过了六年愉快的时光,先后获得人工智能学士学位(一等荣誉学位)和机器人学硕士学位(Distinction)。
我的研究致力于构建能够在开放场景中理解、生成与泛化的多模态模型:
- 多模态大语言模型 – 学术图表、表格与文档的细粒度理解
- 语义对齐 – 视觉、音频与语言等异构模态之间的语义对齐
- 开放场景泛化 – 多模态大模型向未见视觉模态(热红外、深度、X 光等)的零样本迁移
- 3D 场景生成 – 基于结构化布局先验的可控文本到 3D 场景生成
多模态大模型 语义对齐 开放场景泛化 3D 场景生成 机器人
非常欢迎交流科研与合作!可以通过 3242701514@qq.com 联系我,或者在 GitHub 上找到我。
# 近期动态
- 2026 – 两篇论文被 ECCV 2026 接收:VVM-Tuning(多模态大模型的多视觉模态泛化)与 ScrollScape(基于视频扩散先验的 32K 图像生成)。
- 2026 – 第一作者论文 ControlScene 被 CAAI Transactions 接收。
- 2026 – 任务不确定性感知视频复原论文被 IEEE TMM 接收。
- 2025.03 – 进入哈尔滨工业大学攻读博士学位。
- 2024.12 – 以 Distinction 成绩获得曼彻斯特大学机器人学硕士学位。
# 教育背景
计算机科学与技术(博士在读)· 哈尔滨工业大学 · 2025.03 -- 至今
- 导师:左旺孟教授
- 研究方向:多模态大语言模型、语义对齐、开放场景泛化、3D 场景生成
机器人学(硕士)· 曼彻斯特大学 · 2022.09 -- 2024.12
- 学位:一等荣誉学位(Distinction)
- 毕业论文:水下机器人视觉感知与扩散模型先验建模
人工智能(学士)· 曼彻斯特大学 · 2018.09 -- 2022.06
- 学位:一等荣誉学位(First Class Honours)
- 荣誉:一等本科毕业论文(Best Undergraduate Dissertation)
# 研究经历
多模态智能实验室,哈尔滨工业大学・博士研究员・2025.03 – 至今
- 主导学术文献多模态理解方向,设计 Count-Curriculum 课程学习框架,解决困难样本梯度弱的问题,将图表匹配准确率提升 23.3%。
- 提出结构化布局引导的 3D 场景生成方法 ControlScene,构建包含 20,000 个人工标注样本的 LayoutVerse 基准。
- 探索开放场景下多模态大模型的语义对齐机制,形成 "理解 – 生成 – 泛化" 的研究体系。
机器人实验室,曼彻斯特大学・硕士研究员・2022.09 – 2024.12
- 研究水下机器人视觉感知,开发基于扩散模型的图像先验建模方法,显著提升低光照与浑浊环境下的检测鲁棒性。
- 设计并优化区域兴趣视觉(RoIV)模型,结合外部注意力机制,在计算资源受限条件下实现实时目标检测。
生物机器学习项目,麻省理工学院(MIT) · 研究助理(远程) · 2021.06 – 2021.09
- 作为团队中年龄最小的成员参与生物序列数据的深度学习建模,负责数据预处理与特征工程。
- 协助设计用于 scRNA 与 scATAC 数据翻译的多任务自编码器框架,项目获组内排名第一。
黑龙江哈工智能产业有限公司・算法工程师(实习) · 2022.06 – 2022.09
- 优化用于图像识别的 CNN 架构,进行超参数调整与模型蒸馏,提升识别准确率。
- 微调 NLP 模型用于文本语义理解,支持工业场景下的智能数据分析。
# 精选项目
- 自主机器人设计与组装(硕士核心项目,2023.09 – 2024.02)-- 设计并组装自主机器人,集成路径规划与计算机视觉算法,实现复杂环境中的导航与机械臂精准操控;优化控制系统延迟,提高动态环境下的响应效率与任务成功率。
- 高维数据降维方法比较分析(本科毕业论文)-- 系统评估 PCA、LDA、t-SNE 与 UMAP 在噪声数据上的表现,发现 UMAP 在保持局部结构方面的优势,获评一等本科毕业论文。
- 六角棋(Hex)AI 博弈算法(2022.11)-- 合作开发基于 FastVC(虚拟连接)搜索的评价策略,融合启发式规则与开局交换策略,在对阵其他同学开发的 AI 时取得约 75% 的胜率。
- 更多课程项目请见首页 – 国际象棋、迷宫求解、贪吃蛇、模糊逻辑洗衣机等。
# 荣誉与奖项
- 硕士一等学位(Distinction)、本科一等学位(First Class)、一等本科毕业论文
- 2022 华为软件精英挑战赛 – 全国第 71 名(前 3%)
- 2020 公益网站设计大赛 – 最佳创意奖
- 曼彻斯特大学 PASS Leader;本科优秀学生代表
# 技术技能
–|--
编程语言 | Python, PyTorch, TensorFlow, C++, Bash, SQL, LaTeX
多模态 / 大模型 | Qwen-VL, LLaVA, ImageBind, CLIP, BLIP-2, LoRA / QLoRA
3D / 图形学 | 3D Gaussian Splatting, NeRF, Open3D, PyTorch3D
工具平台 | Linux, Git, Docker, Slurm, Weights & Biases, OpenCV
语言能力 | 中文(母语),英文(流利,雅思 7.0,英国本硕全英文授课)