Research direction

AI for Database Optimisation

This research programme has developed through four stages. Lemo replaces the traditional query optimiser with a learned optimiser for concurrent workloads, while RankPQO enhances the traditional query optimiser for parametric workloads by adding learning-to-rank. TATA transfers plan representations between database tasks, and ScalePQO brings shared, adaptive models into OceanBase.

Research progression

  1. Replacement

    Lemo

    Learned latency estimation and reuse-aware plan generation for concurrent queries.

  2. Enhancement

    RankPQO

    Adds learning-to-rank to the traditional query optimiser for parametric workloads.

  3. Transfer

    TATA

    Query-plan representations transfer across database tasks.

  4. Industrial deployment

    ScalePQO

    Shared, adaptive models operate inside OceanBase.

Concurrent queries · ACM SIGMOD 2024

Lemo

Concurrent queries can repeat the same computation while competing for shared resources. Lemo combines learned latency estimation with a shared buffer for intermediate subquery results, so cached results can be considered during plan generation. The system was implemented as a PostgreSQL prototype.

I am the first author and implemented the PostgreSQL prototype.

Parametric plans · VLDB 2025

RankPQO

Parameterized SQL templates may require different plans for different parameter bindings, while absolute latency is difficult to predict reliably. RankPQO combines plan enumeration with learning-to-rank for offline plan caching and online plan selection in PostgreSQL.

I am the first author and led the design and implementation of the query-optimisation work.

The paper reports up to 2.57× acceleration over the PostgreSQL optimiser and up to 1.36× over the strongest evaluated baseline.

Task transfer · VLDB 2026

TATA

Learned database components often require new labelled data for every task. TATA transfers query-plan representations across tasks by combining self-supervised plan reconstruction with database-aware generation of pseudo-labelled target data. It was evaluated when transferring from cost estimation to query optimisation and index selection.

I am a contact author of this work.

The paper reports up to a fivefold reduction in data-collection cost across the evaluated transfer settings.

Industrial system · VLDB 2026 Industry Track

ScalePQO

Industrial parametric query optimisation must support many query templates and adapt when parameter distributions change. ScalePQO shares models across related templates and controls online model updates as workloads evolve. The system was implemented in OceanBase, an enterprise-grade distributed database originating from Ant Group.

I am the first author and led the technical development of the query-optimisation strand in the NTU–OceanBase collaboration.

Across six evaluated workloads, the paper reports up to 1.62× acceleration over OceanBase and up to 1.23× over RankPQO.