From query optimisation to complete data systems

My research asks where learned components can improve a database system, and which database guarantees the surrounding system must retain. The work spans algorithms, database internals, and end-to-end implementation.

Current direction

Semantic Query Processing

Semantic workloads place expensive model-backed operators in the same plans as conventional relational operators. I am studying how a database system should execute and optimise these mixed workloads while retaining control over latency, resource use, and result quality.

Status: ongoing research. Project names, manuscript details, and experimental results will be added when they are public.

System building · ICDE 2026

Heterogeneous and Multimodal Data Systems

ARCADE is an open-source system for real-time hybrid and continuous queries over vector, spatial, text, image, and relational data, built on RocksDB and MySQL. As co-first author, I led the design and implementation of its query optimisation and incremental processing for continuous queries; collaborators held primary ownership of storage and indexing.

Evidence: IEEE ICDE 2026; open-source implementation. Project detailsPaperCode

Research foundations · SIGMOD and PVLDB

AI for Database Optimisation

This work studies learned search for concurrent queries, learning-to-rank for parametric plan selection, transferable query-plan representations, and industrial-scale optimisation. It includes PostgreSQL prototypes for Lemo and RankPQO, the TATA transfer framework, and ScalePQO, which was integrated into OceanBase.

Outputs: ACM SIGMOD 2024 and PVLDB / VLDB 2025–2026. Project details

Earlier research

Algorithms for Data-Intensive Applications

My earlier work developed algorithms for graph analytics and urban data applications. It includes weighted random-walk domination, public-transport scheduling, and outdoor-advertising placement over large trajectory and graph datasets.

These projects provided the algorithmic foundation for my later systems work: formal problem definition, approximation and pruning techniques, and empirical evaluation on real datasets.

Selected outputs: PVLDB 2021, IEEE TKDE, and ACM KDD 2019 Best Paper Runner-up.

The publication record behind the programme

View all publications