Selected Research Projects

Research theme

Brain connectivity & statistical neuroimaging

Orthogonal sparse subgraph alignment
OSSA pipeline: construct functional communities, learn a within-community orthogonal alignment with sparse node selection, and recover compact structural subgraphs concordant with function.

Orthogonal Sparse Subgraph Alignment for Structure-Function Coupling in Brain Networks

2026 | NeurIPS Under review

A principled statistical framework that localizes cross-modal subnetworks in structure–function (SC–FC) coupling by jointly learning an orthogonal transformation and a sparse node-selection vector. It generalizes orthogonal Procrustes to an unknown induced-subgraph target, comes with consistency guarantees, and achieves state-of-the-art SC–FC alignment on the ABCD study, outperforming nine size-matched and deep-learning baselines.

Research theme

Probabilistic modeling for multimorbidity in longitudinal health data

Bayesian hypergraph pathway inference
Bayesian hypergraph workflow linking patient covariates and disease outcomes to latent disease clusters and shared risk pathways.

Disentangling Latent Risk Pathways via Bayesian Hypergraph Inference

2026 | ICML, Oral (0.7% acceptance)

A Bayesian hypergraph framework for longitudinal EHR that jointly models multiple disease onsets and survival, discovers risk-factor-specific higher-order disease clusters, and scales to large biobank data through structured mean field variational inference.

Research theme

Multilingual NLP & LLM evaluation

OasisSimp dataset
Complex sentences paired with expert simplifications across English, Sinhala, Tamil, Thai, and Pashto, annotated with edit operations (deletion, rewording).

OasisSimp: An Open-source Asian–English Sentence Simplification Dataset

2026 | LREC, Oral

An open-source sentence-simplification dataset for five languages (English, Sinhala, Tamil, Thai, Pashto) with expert human simplifications, benchmarking eight open-weight multilingual LLMs in zero-/few-shot settings to expose low-resource simplification gaps.

MultiMWP corpus
MultiMWP generation pipeline: token embeddings drive an MWP generator and equation generator, trained under an equation-consistency constraint and reconstruction loss.

A Multilingual Dataset (MultiMWP) and Benchmark for Math Word Problem Generation

2025 | IEEE/ACM TASLP (Vol. 33)

The largest multi-way parallel math-word-problem corpus to date, spanning nine languages (six low-resource), benchmarking pretrained multilingual seq2seq models (mBART50, mT5, M2M-100, IndicBART) with a math-constraint generation module.

SIB-200 benchmark
Topic-classification accuracy across 200+ languages and dialects for several multilingual models, exposing the gap between high-resource and unseen low-resource languages.

SIB-200: A Simple, Inclusive, and Big Evaluation Dataset for Topic Classification in 200+ Languages and Dialects

2024 | EACL

A topic-classification benchmark covering 205 languages and dialects (the first NLU evaluation set for many of them), built on the Flores-200 parallel corpus, with evaluation of multilingual PLMs under supervised, cross-lingual transfer, and LLM-prompting settings.

Research theme

Financial GenAI & large-scale retrieval

Financial LLM-RAG pipeline
LLM–RAG pipeline: text processing and embedding, multi-stage retrieval with no-look-ahead constraints, evidence compression, and constrained project generation with optional QLoRA calibration.

End-to-end LLM–RAG Pipeline for Financial Estimates

Ongoing

An end-to-end LLM–RAG system over 2TB+ of heterogeneous financial text (SEC 10-K filings, USPTO patents, earnings-call transcripts, WSJ articles) with multi-stage semantic retrieval, strict no-look-ahead temporal constraints, and high-throughput batched serving — improving project-identification accuracy over a keyword-retrieval baseline.

Industry

Machine Learning Engineer — DeepManifold Inc., New York

Part-time Intern · Sep 2025 – Dec 2025

Built a customer embedding platform using sequence models (Word2Vec, Transformer) and contrastive self-supervised learning (SimCLR, MoCo) for 50M+ users, and operationalized segmentation workflows (KMeans, GMM, DBSCAN, extended RFM) — personalized campaigns increased CTR by 14% and reduced churn by 9% in controlled experiments.

Cloud Data Engineer — Huawei Technologies Canada, Toronto

Full-time · Apr 2022 – May 2023

Developed a distributed, low-latency data engine and an InfluxDB connector for cross-data-center migration across 40+ databases — now deployed at top-5 banks in China serving 3M+ users — and implemented SQL predicate push-down that improved TPC-H 1000GB query speed by an average of 20%.