Research Threads: Scaling Data Intelligence


From sequence alignment to user understanding, from Burrows–Wheeler transform to Graph RAG, the apparent diversity of topics resolves into one through-line: making large-scale data searchable, matchable, and understandable. The domains changed — genomes, then user and transaction graphs, now knowledge and language — and one methodology persisted throughout. Proposed around 2015–2016 and named Compact Computing (凝练计算), it is the foundation beneath every thread on this page rather than a thread of its own: a style of computing that centers on data and fuses algorithms with systems for full-stack performance. Its signature is operating directly on compressed, sparse, or elastic representations — Burrows–Wheeler transform and suffix arrays for genome-scale indexing, CSR-based sparse matrix–vector multiplication on GPUs, compiler-generated sparse kernels for GNNs — and it now extends to the foundation-model era through KV-cache reuse for RAG and tree speculation for hybrid-attention language models. Compact Computing consists of three components in principle: (1) tightly-coupled architectures (CPUs, GPUs, Xeon Phis, clusters), (2) compressive and elastic data representation, and (3) efficient, scalable, service-oriented algorithms. This page reorganizes the work along five threads; the full publication list, with every entry tagged by its thread, is on the Publications page.

T1 · Sequence Intelligence

The first act, 2009–2017: in retrospect, a search engine for the genome. The pipeline was always the same one we now use everywhere else — build a compact index (BWT, suffix array, k-mer tables), search against it with approximate matching (Smith–Waterman, seed–extend alignment), and understand the results (variant calling, metagenomic quantification). The CUSHAW suite, CUDASW++ and SWAPHI families of tools were accelerated across the full hardware spectrum of the day — CUDA GPUs, Xeon Phi coprocessors, UPC++ clusters — and three of them were rated by NVIDIA as popular GPU-accelerated applications; DecGPU was reported by GenomeWeb. This era internalized the habits that still shape the work: domain problem → compact data structure → parallel algorithm → open-source tool.

Example work

Several of this era's open-source tools (CUDASW++, CUSHAW, MSAProbs, Musket, SWAPHI, ParaBWT, LightPCC and more) are catalogued on the Software page.

T2 · Graph Intelligence

The second act began at Ant Group in 2017, when the "database being searched" became the graph: hundred-million-user banking relationships, transactions, and behaviors. The graph itself had in fact entered the picture during the genome years: de Bruijn graph assembly of large genomes (2011) was the first large-graph compute problem tackled, and LightSpMV (2015, Best Paper Award) distilled the workhorse primitive of graph analytics — CSR-based sparse matrix–vector multiplication on GPUs — into a general-purpose kernel. At Ant, he led GeaLearning (alias GraphTheta) — the first distributed, scalable graph learning system built on the vertex-centric graph processing paradigm — whose landing on Zhima Credit in Alipay was listed in Ant Technology Memorabilia 2020 and which is a core component of GeaGraph (alias TuGraph), recognized as a World's Leading Internet Scientific and Technological Achievement at the World Internet Conference 2021. The team broke the world record on the Stanford Open Graph Benchmark proteins leaderboard in 2021. Beyond the flagship system, the thread covers the temporal-graph stack (TeGraph, TEA, TeMatch), compiler and GPU optimizations for GNN training/inference, and graph learning methodology from heterogeneous graph architecture search, heterogeneous graph augmentation, and materialization-free community detection (SIGMOD 2025) to graph attention designs, benchmarks and defenses for text-attributed graph learning, and tabular foundation models as graph anomaly detectors; the foundation-model line continues in Thread T3.

Example work

T3 · Foundation Models: from Graphs to Users

GraphCLIP and the FIND/FOUND series are one technology applied to two worlds: pretrain over large graphs, then transfer across domains, scenarios and tasks. GraphCLIP brings the playbook to text-attributed graphs, making graph foundation models transferable across domains; the FIND series brings it to user graphs — from the first transferable and forecastable user targeting foundation model, through FOUNDv2's unified user quantized tokenizers for user representation, to Query-as-Anchor's scenario-adaptive user representation via a large language model — turning user understanding into a pretrain-once, transfer-everywhere capability.

Example work

T4 · User Understanding & Trust Intelligence

Graph intelligence points at people. This thread turns the engines of T3 into understanding of users and protection of trust: recommendation as the matching of users to items over big graphs — from path-based candidate item matching in recommenders, CTR prediction, multi-behavior sequential and cold-start recommendation to agentic point-of-interest recommendation; and trust intelligence as graph-powered risk control for industrial transaction networks, backbone of the award-winning Alipay credit system and of the joint industry white paper 图风控行业技术报告. The user-foundation-model line of this story sits in Thread T3.

Example work

T5 · Retrieval-Augmented Reasoning & Agents

The current act: when the data became language and knowledge, the through-line — index, retrieve, match, understand — re-emerged as RAG. He is the lead author of the globally recognized survey "Graph Retrieval-Augmented Generation: A Survey" (ACM TOIS 2026, ESI Highly Cited Paper), and the thread extends it in both directions: reasoning quality — graph-trajectory-augmented reinforcement learning for multi-turn RAG (GTA-RAG), multi-hop multi-setting graph question answering (M3GQA), subgraph retrieval enhanced by graph–text alignment, vision-guided knowledge-graph reasoning with LLMs (Foresight-on-Graph), and multi-agent graph reasoning; and serving efficiency — declarative auto-optimization of RAG pipelines (AutoRAGTuner), KV-cache reuse (Decoupled Attention Fusion), long-context prefilling acceleration (CutAttn) and tree speculation for hybrid-attention models (Bole) — the Compact Computing habits applied to LLMs.

Example work

Explorations Beyond the Threads


A few pieces of work do not sit inside the five threads but enrich them: federated and privacy-preserving learning (DiVerFed, Best Student Paper at KSEM 2024; "Integer is Enough"), general optimization methodology for deep learning (AGD at NeurIPS 2023; weighted sharpness-aware minimization at KDD 2023), hardware-aware acceleration of deep learning (Woodpecker-DL), and earlier heterogeneous-computing collaborations such as the PyCAC concurrent atomistic-continuum simulation environment.