From query optimisation to complete data systems
My research asks where learned components can improve a database system, and which database guarantees the surrounding system must retain. The work spans algorithms, database internals, and end-to-end implementation.
Current direction
Semantic Query Processing
Semantic workloads place expensive model-backed operators in the same plans as conventional relational operators. I am studying how a database system should execute and optimise these mixed workloads while retaining control over latency, resource use, and result quality.
Status: ongoing research. Project names, manuscript details, and experimental results will be added when they are public.
System building · ICDE 2026
Heterogeneous and Multimodal Data Systems
ARCADE is an open-source system for real-time hybrid and continuous queries over vector, spatial, text, image, and relational data, built on RocksDB and MySQL. As co-first author, I led the design and implementation of its query optimisation and incremental processing for continuous queries; collaborators held primary ownership of storage and indexing.
Evidence: IEEE ICDE 2026; open-source implementation. Project detailsPaperCode
Research foundations · SIGMOD and PVLDB
AI for Database Optimisation
This work studies learned search for concurrent queries, learning-to-rank for parametric plan selection, transferable query-plan representations, and industrial-scale optimisation. It includes PostgreSQL prototypes for Lemo and RankPQO, the TATA transfer framework, and ScalePQO, which was integrated into OceanBase.
Outputs: ACM SIGMOD 2024 and PVLDB / VLDB 2025–2026. Project details
Earlier research
Algorithms for Data-Intensive Applications
My earlier work developed algorithms for graph analytics and urban data applications. It includes weighted random-walk domination, public-transport scheduling, and outdoor-advertising placement over large trajectory and graph datasets.
These projects provided the algorithmic foundation for my later systems work: formal problem definition, approximation and pruning techniques, and empirical evaluation on real datasets.
Selected outputs: PVLDB 2021, IEEE TKDE, and ACM KDD 2019 Best Paper Runner-up.