# Hiiiiiii, I am Ruyi (Maggie) Zhang

I am a PhD student in Computer Science and Technology at the Harbin Institute of Technology (HIT), supervised by Prof. Wangmeng Zuo. Before that, I spent six happy years at the University of Manchester, where I received a BSc in Artificial Intelligence (First Class Honours) and an MSc in Robotics (Distinction).

My research focuses on building multimodal models that understand, generate and generalize in open-world scenarios:

  • Multimodal Large Language Models – fine-grained understanding of academic figures, tables and documents
  • Semantic Alignment – aligning vision, audio and language across heterogeneous modalities
  • Open-Scene Generalization – zero-shot transfer of LMMs to unseen visual modalities (thermal, depth, X-ray, …)
  • 3D Scene Generation – controllable text-to-3D scene generation with structured layout priors

Multimodal LLM Semantic Alignment Open-World Generalization 3D Scene Generation Robotics

I am always happy to chat about research and collaboration. Feel free to reach me at 3242701514@qq.com or find me on GitHub.

# News

  • 2026 – Two papers accepted to ECCV 2026: VVM-Tuning (generalizing LMMs to versatile visual modalities) and ScrollScape (32K image generation with video diffusion priors).
  • 2026 – First-author paper ControlScene accepted to CAAI Transactions.
  • 2026 – Paper on task-uncertainty-aware video restoration accepted to IEEE TMM.
  • 2025.03 – Started my PhD at the Harbin Institute of Technology.
  • 2024.12 – Graduated from the University of Manchester with an MSc in Robotics (Distinction).

# Education

Ph.D. in Computer Science and Technology · Harbin Institute of Technology · 2025.03 -- Present
  • Supervisor: Prof. Wangmeng Zuo
  • Research: multimodal large language models, semantic alignment, open-scene generalization, 3D scene generation
M.Sc. in Robotics · University of Manchester · 2022.09 -- 2024.12
  • Degree: Distinction
  • Dissertation: Visual perception for underwater robots with diffusion-model image priors
B.Sc. in Artificial Intelligence · University of Manchester · 2018.09 -- 2022.06
  • Degree: First Class Honours
  • Honour: Best Undergraduate Dissertation

# Research Experience

Multimodal Intelligence Lab, Harbin Institute of Technology · PhD Researcher · 2025.03 – Present

  • Leading research on multimodal understanding of academic literature. Designed the Count-Curriculum framework, which tackles weak gradients on hard samples and raises figure-matching accuracy by 23.3%.
  • Proposed ControlScene, a structured-layout-guided 3D scene generation method, and built the LayoutVerse benchmark with 20,000 manually annotated samples.
  • Exploring semantic alignment mechanisms of multimodal large models in open scenarios, forming an “understanding – generation – generalization” research line.

Robotics Lab, University of Manchester · MSc Researcher · 2022.09 – 2024.12

  • Studied visual perception for underwater robots; developed diffusion-model-based image prior modeling that notably improves detection robustness in low-light and turbid water.
  • Designed and optimized a Region-of-Interest Vision (RoIV) model with external attention for real-time object detection under limited compute.

Biological Machine Learning Project, MIT · Research Assistant (Remote) · 2021.06 – 2021.09

  • The youngest member of the team, responsible for data preprocessing and feature engineering for deep learning on biological sequence data.
  • Helped design a multi-task autoencoder for translating between scRNA and scATAC data; the project ranked first in the group.

Heilongjiang HIT AI Industry Co., Ltd. · Algorithm Engineer Intern · 2022.06 – 2022.09

  • Optimized CNN architectures for image recognition via hyper-parameter tuning and model distillation.
  • Fine-tuned NLP models for textual semantic understanding in industrial data analysis.

# Selected Projects

  • Autonomous Robot Design & Assembly (MSc core project, 2023.09 – 2024.02) – Designed and built an autonomous robot integrating path planning and computer vision for navigation and precise robotic-arm manipulation in complex environments; reduced control latency to improve responsiveness in dynamic scenes.
  • Comparative Analysis of Dimensionality Reduction (BSc dissertation) – Systematically evaluated PCA, LDA, t-SNE and UMAP on noisy high-dimensional data and showed UMAP’s advantage in preserving local structure. Awarded Best Undergraduate Dissertation.
  • Hex AI Bot (2022.11) – Co-developed an evaluation strategy based on FastVC (virtual connection) search with heuristic rules and a swap-opening strategy, reaching about 75% win rate against other AI agents.
  • More course projects are on the home page – chess, maze solver, snake, fuzzy-logic washing machine and more.

# Honours & Awards

  • MSc with Distinction; BSc with First Class Honours; Best Undergraduate Dissertation
  • Huawei Software Elite Challenge 2022 – National rank 71 (top 3%)
  • 2020 Public Welfare Website Design Competition – Best Creativity Award
  • PASS Leader, University of Manchester; Outstanding Undergraduate Student Representative

# Skills

–|--
Programming | Python, PyTorch, TensorFlow, C++, Bash, SQL, LaTeX
Multimodal / LLM | Qwen-VL, LLaVA, ImageBind, CLIP, BLIP-2, LoRA / QLoRA
3D / Graphics | 3D Gaussian Splatting, NeRF, Open3D, PyTorch3D
Tools | Linux, Git, Docker, Slurm, Weights & Biases, OpenCV
Languages | Chinese (native), English (fluent, IELTS 7.0, BSc & MSc taught in English)

# 你好呀,我是张茹怡 (Maggie)

我目前在哈尔滨工业大学计算机科学与技术专业攻读博士学位,导师是左旺孟教授。在此之前,我在曼彻斯特大学度过了六年愉快的时光,先后获得人工智能学士学位(一等荣誉学位)和机器人学硕士学位(Distinction)。

我的研究致力于构建能够在开放场景中理解、生成与泛化的多模态模型:

  • 多模态大语言模型 – 学术图表、表格与文档的细粒度理解
  • 语义对齐 – 视觉、音频与语言等异构模态之间的语义对齐
  • 开放场景泛化 – 多模态大模型向未见视觉模态(热红外、深度、X 光等)的零样本迁移
  • 3D 场景生成 – 基于结构化布局先验的可控文本到 3D 场景生成

多模态大模型 语义对齐 开放场景泛化 3D 场景生成 机器人

非常欢迎交流科研与合作!可以通过 3242701514@qq.com 联系我,或者在 GitHub 上找到我。

# 近期动态

  • 2026 – 两篇论文被 ECCV 2026 接收:VVM-Tuning(多模态大模型的多视觉模态泛化)与 ScrollScape(基于视频扩散先验的 32K 图像生成)。
  • 2026 – 第一作者论文 ControlScene 被 CAAI Transactions 接收。
  • 2026 – 任务不确定性感知视频复原论文被 IEEE TMM 接收。
  • 2025.03 – 进入哈尔滨工业大学攻读博士学位。
  • 2024.12 – 以 Distinction 成绩获得曼彻斯特大学机器人学硕士学位。

# 教育背景

计算机科学与技术(博士在读)· 哈尔滨工业大学 · 2025.03 -- 至今
  • 导师:左旺孟教授
  • 研究方向:多模态大语言模型、语义对齐、开放场景泛化、3D 场景生成
机器人学(硕士)· 曼彻斯特大学 · 2022.09 -- 2024.12
  • 学位:一等荣誉学位(Distinction)
  • 毕业论文:水下机器人视觉感知与扩散模型先验建模
人工智能(学士)· 曼彻斯特大学 · 2018.09 -- 2022.06
  • 学位:一等荣誉学位(First Class Honours)
  • 荣誉:一等本科毕业论文(Best Undergraduate Dissertation)

# 研究经历

多模态智能实验室,哈尔滨工业大学・博士研究员・2025.03 – 至今

  • 主导学术文献多模态理解方向,设计 Count-Curriculum 课程学习框架,解决困难样本梯度弱的问题,将图表匹配准确率提升 23.3%。
  • 提出结构化布局引导的 3D 场景生成方法 ControlScene,构建包含 20,000 个人工标注样本的 LayoutVerse 基准。
  • 探索开放场景下多模态大模型的语义对齐机制,形成 "理解 – 生成 – 泛化" 的研究体系。

机器人实验室,曼彻斯特大学・硕士研究员・2022.09 – 2024.12

  • 研究水下机器人视觉感知,开发基于扩散模型的图像先验建模方法,显著提升低光照与浑浊环境下的检测鲁棒性。
  • 设计并优化区域兴趣视觉(RoIV)模型,结合外部注意力机制,在计算资源受限条件下实现实时目标检测。

生物机器学习项目,麻省理工学院(MIT) · 研究助理(远程) · 2021.06 – 2021.09

  • 作为团队中年龄最小的成员参与生物序列数据的深度学习建模,负责数据预处理与特征工程。
  • 协助设计用于 scRNA 与 scATAC 数据翻译的多任务自编码器框架,项目获组内排名第一。

黑龙江哈工智能产业有限公司・算法工程师(实习) · 2022.06 – 2022.09

  • 优化用于图像识别的 CNN 架构,进行超参数调整与模型蒸馏,提升识别准确率。
  • 微调 NLP 模型用于文本语义理解,支持工业场景下的智能数据分析。

# 精选项目

  • 自主机器人设计与组装(硕士核心项目,2023.09 – 2024.02)-- 设计并组装自主机器人,集成路径规划与计算机视觉算法,实现复杂环境中的导航与机械臂精准操控;优化控制系统延迟,提高动态环境下的响应效率与任务成功率。
  • 高维数据降维方法比较分析(本科毕业论文)-- 系统评估 PCA、LDA、t-SNE 与 UMAP 在噪声数据上的表现,发现 UMAP 在保持局部结构方面的优势,获评一等本科毕业论文。
  • 六角棋(Hex)AI 博弈算法(2022.11)-- 合作开发基于 FastVC(虚拟连接)搜索的评价策略,融合启发式规则与开局交换策略,在对阵其他同学开发的 AI 时取得约 75% 的胜率。
  • 更多课程项目请见首页 – 国际象棋、迷宫求解、贪吃蛇、模糊逻辑洗衣机等。

# 荣誉与奖项

  • 硕士一等学位(Distinction)、本科一等学位(First Class)、一等本科毕业论文
  • 2022 华为软件精英挑战赛 – 全国第 71 名(前 3%)
  • 2020 公益网站设计大赛 – 最佳创意奖
  • 曼彻斯特大学 PASS Leader;本科优秀学生代表

# 技术技能

–|--
编程语言 | Python, PyTorch, TensorFlow, C++, Bash, SQL, LaTeX
多模态 / 大模型 | Qwen-VL, LLaVA, ImageBind, CLIP, BLIP-2, LoRA / QLoRA
3D / 图形学 | 3D Gaussian Splatting, NeRF, Open3D, PyTorch3D
工具平台 | Linux, Git, Docker, Slurm, Weights & Biases, OpenCV
语言能力 | 中文(母语),英文(流利,雅思 7.0,英国本硕全英文授课)

Give me a cup of [coffee]~( ̄▽ ̄)~*请我喝[茶]~( ̄▽ ̄)~*

Ruyi Zhang WeChat Pay

WeChat Pay

Ruyi Zhang Alipay

Alipay

Ruyi Zhang PayPal

PayPal