Senior Staff Research Scientist (Tech Lead) at Bytedance, USA

I obtained PhD from University of California, Los Angeleunder the supervision of Prof. Alan Yuille, and received M.S/B.S from Peking University . My research interests lie in unifying multi-task models into a single model for generation and understanding. I worked mostly on Computer Vision and Machine Learning, such as learning 2D/3D representations, learning neural architectures, and mining informative instances from large image/video datasets.  

At Bytedance, we delivered techs to lots of Product such as Doubao剪映 CapCut抖音/TikTok .

Follow Me On:

Products

READ MORE
Selected products, where I lead the effort in algorithm design, development and implementation (linked to some demo videos)

SeedEdit 3.0 for general image editing (>50M DAU)

SeedEdit 3.0, which significantly improves over our previous SeedEdit1.0 versions in both aspects of edit instruction following and image content (e.g., ID/IP) preservation on real image inputs. Additional to model upgrading with T2I, in this report, we present several key improvements. First, we develop an enhanced data curation pipeline with a meta-info paradigm and meta-info embedding strategy that help mix images from multiple data sources. This allows us to scale editing data effectively, and meta information is helpfult to connect VLM with diffusion model more closely. Second, we introduce a joint learning pipeline for computing a diffusion loss and reward losses. 

Seedream3.0

Seedream 3.0 is a high-performance Chinese-English bilingual image generation foundation model. We develop several technical improvements to address existing challenges in Seedream 2.0, including alignment with complicated prompts, fine-grained typography generation, suboptimal visual aesthetics and fidelity, and limited image resolutions. Specifically, the advancements of Seedream 3.0 stem from improvements across the entire pipeline, from data construction to model deployment. At the data stratum, we double the dataset using a defect-aware training paradigm and a dual-axis collaborative data-sampling framework. Furthermore, we adopt several effective techniques such as mixed-resolution training, cross-modality RoPE, representation alignment loss, and resolution-aware timestep sampling in the pre-training phase. During the post-training stage, we utilize diversified aesthetic captions in SFT, and a VLM-based reward model with scaling, thereby achieving outputs that well align with human preferences. Furthermore, Seedream 3.0 pioneers a novel acceleration paradigm. By employing consistent noise expectation and importance-aware timestep sampling, we achieve a 4 to 8 times speedup while maintaining image quality. Seedream 3.0 demonstrates significant improvements over Seedream 2.0: it enhances overall capabilities, in particular for text-rendering in complicated Chinese characters which is important to professional typography generation. In addition, it provides native high-resolution output (up to 2K), allowing it to generate images with high visual quality.

3D Photo Zoom Effect [3m DAU increase, help capcut rank #1 in ios free app]

It is well known that deep neural networks are universal function approximators. We adopt a network for human segmentation and depth estimation. 


Cyberpunk photo effect [Top effect in multi-countries e.g. JP/KR]

A technique that connects mobile SLAM, depth/normal estimation, and box detection for virtual effect development. 

AR City Effect [5m Users in Douyin(China Tiktok)]

Talk at NeurIPS robotics learning workshop by Hao Su, Kaichun Mo, and Fanbo Xiang.

It is well known that deep neural networks are universal function approximators and have good generalizability when the training and test datasets are sampled from the same distribution. Most deep learning-based applications and theories in the past decade are based upon this setup. While the view of learning function approximators has been rewarding to the community, we are seeing more and more of its limitations when dealing with the real-world problem space that is combinatorially exploded. In this talk, I will discuss a possible shift of view, from learning function approximators to learning algorithm approximators, by some preliminary work in my lab. Our ultimate goal is to achieve generalizability when learning in a problem space of combinatorial complexity. We refer to this desired generalizability as compositional generalizability.

Virtual Object AR Attachment [1st 3D virtual effects in TT]

Talk at NeurIPS robotics learning workshop by Hao Su, Kaichun Mo, and Fanbo Xiang.

It is well known that deep neural networks are universal function approximators 

SkyAR Effect

A mobile network that does sky segmentation, which provides fancy effects for our users.