Research Threads: Scaling Data Intelligence
From sequence alignment to user understanding, from Burrows–Wheeler transform to Graph RAG, the apparent diversity of topics resolves into one through-line: making large-scale data searchable, matchable, and understandable. The domains changed — genomes, then user and transaction graphs, now knowledge and language — and one methodology persisted throughout. Proposed around 2015–2016 and named Compact Computing (凝练计算), it is the foundation beneath every thread on this page rather than a thread of its own: a style of computing that centers on data and fuses algorithms with systems for full-stack performance. Its signature is operating directly on compressed, sparse, or elastic representations — Burrows–Wheeler transform and suffix arrays for genome-scale indexing, CSR-based sparse matrix–vector multiplication on GPUs, compiler-generated sparse kernels for GNNs — and it now extends to the foundation-model era through KV-cache reuse for RAG and tree speculation for hybrid-attention language models. Compact Computing consists of three components in principle: (1) tightly-coupled architectures (CPUs, GPUs, Xeon Phis, clusters), (2) compressive and elastic data representation, and (3) efficient, scalable, service-oriented algorithms. This page reorganizes the work along five threads; the full publication list, with every entry tagged by its thread, is on the Publications page.
T1 · Sequence Intelligence
The first act, 2009–2017: in retrospect, a search engine for the genome. The pipeline was always the same one we now use everywhere else — build a compact index (BWT, suffix array, k-mer tables), search against it with approximate matching (Smith–Waterman, seed–extend alignment), and understand the results (variant calling, metagenomic quantification). The CUSHAW suite, CUDASW++ and SWAPHI families of tools were accelerated across the full hardware spectrum of the day — CUDA GPUs, Xeon Phi coprocessors, UPC++ clusters — and three of them were rated by NVIDIA as popular GPU-accelerated applications; DecGPU was reported by GenomeWeb. This era internalized the habits that still shape the work: domain problem → compact data structure → parallel algorithm → open-source tool.
Example work
- Yuandong Chan, Kai Xu, Haidong Lan, Weiguo Liu, Yongchao Liu and Bertil Schmidt: "PUNAS: a parallel ungapped-alignment-featured seed verification for next-generation sequencing read alignment". IEEE IPDPS 2017 — the seed–extend alignment line, at its last stop before Ant.
- Haidong Lan, Weiguo Liu, Yongchao Liu and Bertil Schmidt: "SWhybrid: a hybrid parallel framework for large-scale protein sequence database search". IEEE IPDPS 2017 — the SWAPHI family's final continuation, from Xeon Phi to hybrid parallelism.
- Tony Pan, Patrick Flick, Chirag Jain, Yongchao Liu and Srinivas Aluru: "Kmerind: a flexible parallel library for k-mer indexing of biological sequences on distributed memory systems". ACM-BCB 2016 / IEEE/ACM TCBB, 2019.
- Yongchao Liu, Thomas Hankeln, and Bertil Schmidt: "Parallel and space-efficient construction of Burrows-Wheeler transform and suffix array for big genome data". IEEE TCBB, 2016 — the compact index at genome scale.
- Yongchao Liu, Tran TT, Felix Lauenroth and Bertil Schmidt: "SWAPHI-LS: Smith-Waterman algorithm on Xeon Phi coprocessors for long DNA sequences". IEEE Cluster 2014 (Best Paper Award Recommendation).
- Yongchao Liu, Bernt Popp, and Bertil Schmidt: "CUSHAW3: sensitive and accurate base-space and color-space short-read alignment with hybrid seeding". PLOS ONE, 2014 — hybrid seeding over compressed indexes.
- Yongchao Liu, Adrianto Wirawan and Bertil Schmidt: "CUDASW++ 3.0: accelerating Smith-Waterman protein database search by coupling CPU and GPU SIMD instructions". BMC Bioinformatics, 2013.
- Yongchao Liu, Jan Schröder, and Bertil Schmidt: "Musket: a multistage k-mer spectrum based error corrector for Illumina sequence data". Bioinformatics, 2013.
- Yongchao Liu and Bertil Schmidt: "Long read alignment based on maximal exact match seeds". Bioinformatics / ECCB 2012.
- Yongchao Liu, Bertil Schmidt, and Douglas L. Maskell: "CUSHAW: a CUDA compatible short read aligner to large genomes based on the Burrows-Wheeler transform". Bioinformatics, 2012.
- Yongchao Liu, Bertil Schmidt, Weiguo Liu, and Douglas L. Maskell: "CUDA-MEME: accelerating motif discovery in biological sequences using CUDA-enabled graphics processing units". Pattern Recognition Letters, 2010 — the first of the CUDA-accelerated tools.
- Yongchao Liu, Bertil Schmidt, and Douglas L. Maskell: "MSAProbs: multiple sequence alignment based on pair hidden Markov models and partition function posterior probabilities". Bioinformatics, 2010.
Several of this era's open-source tools (CUDASW++, CUSHAW, MSAProbs, Musket, SWAPHI, ParaBWT, LightPCC and more) are catalogued on the Software page.
T2 · Graph Intelligence
The second act began at Ant Group in 2017, when the "database being searched" became the graph: hundred-million-user banking relationships, transactions, and behaviors. The graph itself had in fact entered the picture during the genome years: de Bruijn graph assembly of large genomes (2011) was the first large-graph compute problem tackled, and LightSpMV (2015, Best Paper Award) distilled the workhorse primitive of graph analytics — CSR-based sparse matrix–vector multiplication on GPUs — into a general-purpose kernel. At Ant, he led GeaLearning (alias GraphTheta) — the first distributed, scalable graph learning system built on the vertex-centric graph processing paradigm — whose landing on Zhima Credit in Alipay was listed in Ant Technology Memorabilia 2020 and which is a core component of GeaGraph (alias TuGraph), recognized as a World's Leading Internet Scientific and Technological Achievement at the World Internet Conference 2021. The team broke the world record on the Stanford Open Graph Benchmark proteins leaderboard in 2021. Beyond the flagship system, the thread covers the temporal-graph stack (TeGraph, TEA, TeMatch), compiler and GPU optimizations for GNN training/inference, and graph learning methodology from heterogeneous graph architecture search, heterogeneous graph augmentation, and materialization-free community detection (SIGMOD 2025) to graph attention designs, benchmarks and defenses for text-attributed graph learning, and tabular foundation models as graph anomaly detectors; the foundation-model line continues in Thread T3.
Example work
- Yunhui Liu, Pengyu Qiu, Yu Xing, Peng Du, Yongchao Liu, Chuntao Hong, Jiajun Zheng, Tao Zheng, Tieke He: "Bridging Academia and Industry: a comprehensive benchmark for attributed graph clustering". NeurIPS 2026.
- Yunhui Liu, Yongchao Liu, Yinfeng Chen, Chuntao Hong, Tao Zheng, Tieke He: "Learning hierarchical knowledge in text-rich networks with taxonomy-informed representation learning". KDD 2026.
- Chengying Huan, Zhengyi Yang, Haoshen Yang, Shaonan Ma, Rong Gu, Fang Xi, Yongchao Liu, Guihai Chen, Chen Tian: "Gem: scalable monotonic graph processing beyond billion-scale on a single machine". SIGMOD 2026.
- Yunhui Liu, Tieke He, Yongchao Liu, Can Yi, Hong Jin, Chuntao Hong: "Tabular foundation models are strong graph anomaly detectors". WWW 2026.
- Runlin Lei, Lu Yi, Mingguo He, Pengyu Qiu, Zhewei Wei, Yongchao Liu, Chuntao Hong: "Robustness in text-attributed graph learning: insights, trade-offs, and new defenses". ICLR 2026.
- Jie Peng, Jiarui Ji, Runlin Lei, Zhewei Wei, Yongchao Liu, Chuntao Hong: "GDGB: a benchmark for generative dynamic text-attributed graph learning". ICLR 2026.
- Junyi Mei, Shixuan Sun, Chao Li, Jing Wang, Xiaofeng Hou, Minyi Guo, Yongchao Liu, Chuntao Hong: "DGS: a GPU-based adaptive graph sampling framework". ACM TACO, 2026.
- Xiaotang Wang, Yun Zhu, Haizhou Shi, Yongchao Liu, Chuntao Hong: "Graph triple attention networks: a decoupled perspective". KDD 2025.
- Jiaxin Jiang, Siyuan Yao, Yuhang Chen, Bingsheng He, Yudong Niu, Yuchen Li, Shixuan Sun, Yongchao Liu: "Community detection in heterogeneous information networks without materialization". SIGMOD 2025.
- Chengying Huan, Heng Zhang, Yongchao Liu, Likang Chen, Xuran Wang, Yongchun Jiang, Shaonan Ma, Yanjun Wu: "TeMatch: a fast temporal subgraph matching framework with temporal-aware subgraph matching algorithms". ICDE 2025.
- Yue Jin, Chengying Huan, Heng Zhang, Yongchao Liu, Shuaiwen Leon Song, Rui Zhao, Yao Zhang, Changhua He, Wenguang Chen: "G-Sparse: compiler-driven acceleration for generalized sparse computation for graph neural networks on modern GPUs". PACT 2023.
- Yuchen Zhou, Yanan Cao, Yongchao Liu, Yanmin Shang, Peng Zhang, Zheng Lin, Yun Yue, Baokun Wang, Xing Fu, Weiqiang Wang: "Multi-aspect heterogeneous graph augmentation". WWW 2023.
- Chengying Huan, Shuaiwen Leon Song, Santosh Pandey, Hang Liu, Yongchao Liu, Baptiste Lepers, Changhua He, Kang Chen, Jinlei Jiang, Yongwei Wu: "TEA: a general-purpose temporal graph random walk engine". EuroSys 2023.
- Chengying Huan, Shuaiwen Leon Song, Yongchao Liu, Heng Zhang, Hang Liu, Charles He, Kang Chen, Jinlei Jiang, Yongwei Wu: "T-GCN: a sampling based streaming graph neural network system with hybrid architecture". PACT 2022.
- Chengying Huan, Hang Liu, Mengxing Liu, Yongchao Liu, Changhua He, Kang Chen, Jinlei Jiang, Yongwei Wu, Shuaiwen Leon Song: "TeGraph: a novel general-purpose temporal graph computing engine". ICDE 2022.
- Yongchao Liu, Houyi Li, Guowei Zhang, Xintan Zeng, Yongyong Li, Bin Huang, Peng Zhang, Zhao Li, Xiaowei Zhu, Changhua He, Wenguang Chen: "GraphTheta: a distributed graph neural network learning system with flexible training strategy" (GeaLearning). arXiv:2104.10569, 2021.
- Yongchao Liu and Bertil Schmidt: "LightSpMV: faster CUDA-compatible sparse matrix-vector multiplication using compressed sparse rows". Journal of Signal Processing Systems, 2018, 90(1):69-86.
- Yongchao Liu and Bertil Schmidt: "LightSpMV: faster CSR-based sparse matrix-vector multiplication on CUDA-enabled GPUs". IEEE ASAP 2015 (Best Paper Award) — the kernel beneath vertex-centric graph processing, the paradigm GeaLearning would later build on.
- Yongchao Liu: "OpenGraphAssembly: abstract, modularize and parallelize fundamental building blocks for graph-based genome assembly". ResearchGate, 2015 — the graph view of genome assembly, recast as modular parallel building blocks.
- Yongchao Liu, Bertil Schmidt, and Douglas L. Maskell: "Parallelized short read assembly of large genomes using de Bruijn graphs". BMC Bioinformatics, 2011 — the genome assembled as one giant graph, long before graphs became about people.
T3 · Foundation Models: from Graphs to Users
GraphCLIP and the FIND/FOUND series are one technology applied to two worlds: pretrain over large graphs, then transfer across domains, scenarios and tasks. GraphCLIP brings the playbook to text-attributed graphs, making graph foundation models transferable across domains; the FIND series brings it to user graphs — from the first transferable and forecastable user targeting foundation model, through FOUNDv2's unified user quantized tokenizers for user representation, to Query-as-Anchor's scenario-adaptive user representation via a large language model — turning user understanding into a pretrain-once, transfer-everywhere capability.
Example work
- Chuan He, Yang Chen, Bin Dou, Wuliang Huang, Baokun Wang, Yongchao Liu, Xing Fu, Yu Cheng, Chuntao Hong, Weiqiang Wang, Zhongle Xie, Jiajun Zheng, Xin-Wei Yao: "FOUNDv2: learning unified user quantized tokenizers for user representation". KDD 2026 (oral).
- Jiahao Yuan, Yike Xu, Jinyong Wen, Baokun Wang, Ziyi Gao, Xiaotong Lin, Yun Liu, Xing Fu, Yu Cheng, Yongchao Liu, Weiqiang Wang, Zhongle Xie: "Query as Anchor: scenario-adaptive user representation via large language model". KDD 2026.
- Yun Zhu, Haizhou Shi, Xiaotang Wang, Yongchao Liu, Yaoke Wang, Boci Peng, Chuntao Hong, Siliang Tang: "GraphCLIP: enhancing transferability in graph foundation models for text-attributed graphs". WWW 2025.
- Bin Dou, Baokun Wang, Yun Zhu, Xiaotong Lin, Yike Xu, Xiaorui Huang, Yang Chen, Yun Liu, Shaoshuai Han, Yongchao Liu, Tianyi Zhang, Yu Cheng, Weiqiang Wang and Chuntao Hong: "Transferable and forecastable user targeting foundation model". WWW 2025 (industry track).
T4 · User Understanding & Trust Intelligence
Graph intelligence points at people. This thread turns the engines of T3 into understanding of users and protection of trust: recommendation as the matching of users to items over big graphs — from path-based candidate item matching in recommenders, CTR prediction, multi-behavior sequential and cold-start recommendation to agentic point-of-interest recommendation; and trust intelligence as graph-powered risk control for industrial transaction networks, backbone of the award-winning Alipay credit system and of the joint industry white paper 图风控行业技术报告. The user-foundation-model line of this story sits in Thread T3.
Example work
- Jinze Wang, yangchen zeng, Tiehua Zhang, Lu Zhang, Yuze Liu, Yongchao Liu, Xingjun Ma, Zhu Sun: "Agent4POI: agentic context-conditioned affordance reasoning for multimodal point-of-interest recommendation". ACM MM 2026.
- Yice Luo, Yun Zhu, Xi Chen, Yongchao Liu, Xintan Zeng, Chengying Huan, Kai Zhang, Jinrui Zhang, Juelu Zhang and Jiajun Zheng: "GraphFAS: a distributed system for automated graph feature generation and selection in industrial transaction networks". CIKM 2026 (oral).
- Chuan He, Yongchao Liu, Qiang Li, Chuntao Hong, Wenliang Zhong, Xin-Wei Yao: "M2VAE: multi-modal multi-view variational autoEncoder for cold-start item recommendation". AAAI 2026.
- Chuan He, Yongchao Liu, Qiang Li, Weiqiang Wang, Xing Fu, Xinyi Fu, Chuntao Hong, Xin-Wei Yao: "Multi-grained preference enhanced transformer for multi-behavior sequential recommendation". KDD 2025.
- Sheng Tian, Xintan Zeng, Yifei Hu, Baokun Wang, Yongchao Liu, Yue Jin, Changhua Meng, Chuntao Hong, Tianyi Zhang, Weiqiang Wang: "GraphRPM: risk pattern mining on industrial large attributed graphs". ECML PKDD 2024.
- Yice Luo, Guannan Wang, Yongchao Liu, Jiaxin Yue, Weihong Cheng and Binjie Fei: "FAF: a risk detection framework on industry-scale graphs". CIKM 2023.
- Houyi Li, Zhihong Chen, Chenliang Li, Rong Xiao, Hongbo Deng, Peng Zhang, Yongchao Liu and Haihong Tang: "Path-based deep network for candidate item matching in recommenders". SIGIR 2021.
T5 · Retrieval-Augmented Reasoning & Agents
The current act: when the data became language and knowledge, the through-line — index, retrieve, match, understand — re-emerged as RAG. He is the lead author of the globally recognized survey "Graph Retrieval-Augmented Generation: A Survey" (ACM TOIS 2026, ESI Highly Cited Paper), and the thread extends it in both directions: reasoning quality — graph-trajectory-augmented reinforcement learning for multi-turn RAG (GTA-RAG), multi-hop multi-setting graph question answering (M3GQA), subgraph retrieval enhanced by graph–text alignment, vision-guided knowledge-graph reasoning with LLMs (Foresight-on-Graph), and multi-agent graph reasoning; and serving efficiency — declarative auto-optimization of RAG pipelines (AutoRAGTuner), KV-cache reuse (Decoupled Attention Fusion), long-context prefilling acceleration (CutAttn) and tree speculation for hybrid-attention models (Bole) — the Compact Computing habits applied to LLMs.
Example work
- Wentao Liu, Xiabao Wu, Yongchao Liu, Haitao Zhang, Jiajun Zheng, Ruiting Zhou: "CutAttn: discovering cognitive transition layers for efficient long-context prefilling". NeurIPS 2026.
- Jun Chen, Yongchao Liu, Pengyu Qiu, Jiajun Zheng, Juelu Zhang, Yujie Zeng, Qin Zhang, Xiao Luo, Ziyue Qiao: "GTA-RAG: graph-trajectory-augmented reinforcement learning for multi-turn retrieval-augmented reasoning". EMNLP 2026.
- Yuwei Hu, Runlin Lei, Zhewei Wei, Yongchao Liu, Chuntao Hong, Jiajun Zheng, Juelu Zhang: "Foresight-on-Graph: vision-guided knowledge graph reasoning with large language models". EMNLP 2026.
- Boci Peng, Yun Zhu, Yongchao Liu, Xiaohe Bo, Haizhou Shi, Chuntao Hong, Yan Zhang, Siliang Tang: "Graph retrieval-augmented generation: a survey". ACM TOIS, 2026 (ESI Highly Cited Paper).
- Li Wang, Yi Su, Xiabao Wu, Chiran You, Yongchao Liu, Zhan Qiu, Juelu Zhang, Jiajun Zheng, Fangxin Liu, Jie Zhang, Chen Tian, Chengying Huan: "Bole: efficient tree speculation for hybrid-attention language models". arXiv:2608.01651, 2026.
- Xintan Zeng, Yongchao Liu, Yice Luo, Jiajun Zheng: "AutoRAGTuner: a declarative framework for automatic optimization of RAG pipelines". EuroSys 2026 (poster track).
- Xiabao Wu, Yongchao Liu, Wei Qin, Chuntao Hong: "Decoupled attention fusion: accelerating RAG with efficient KV cache reuse". EuroSys 2026 (poster track).
- Chengying Huan, Ziheng Meng, Yongchao Liu, Zhengyi Yang, Yun Zhu, Yue Yun, Shipeng Li, Rong Gu, Xiabao Wu, Haitao Zhang, Chuntao Hong, Shaonan Ma, Guihai Chen, Chen Tian: "Scaling Graph Chain-of-Thought Reasoning: a multi-agent framework with efficient LLM serving". arXiv:2511.01633, 2025.
- Boci Peng, Yongchao Liu, Xiaohe Bo, Jiaxin Guo, Yun Zhu, Xuanbo Fan, Chuntao Hong, Yan Zhang: "M3GQA: a multi-entity multi-hop multi-setting graph question answering benchmark". ACL 2025.
- Boci Peng, Yongchao Liu, Xiaohe Bo, Sheng Tian, Baokun Wang, Chuntao Hong, Yan Zhang: "Subgraph retrieval enhanced by graph-text alignment for commonsense question answering". ECML PKDD 2024.
Explorations Beyond the Threads
A few pieces of work do not sit inside the five threads but enrich them: federated and privacy-preserving learning (DiVerFed, Best Student Paper at KSEM 2024; "Integer is Enough"), general optimization methodology for deep learning (AGD at NeurIPS 2023; weighted sharpness-aware minimization at KDD 2023), hardware-aware acceleration of deep learning (Woodpecker-DL), and earlier heterogeneous-computing collaborations such as the PyCAC concurrent atomistic-continuum simulation environment.