Retrieval research · Backend & distributed systems
Computer Science undergraduate(B.S. expected 2028) in Taichung, Taiwan.
I study and build systems whose outputs become future state—from retrieved evidence that updates the next query to transaction and failure outcomes that determine the next backend transition. I focus on correctness boundaries, controlled evaluation, and evidence-backed engineering.
- Research: approximate nearest-neighbor retrieval under iterative dense feedback
- Engineering: transactional correctness, durable event delivery, observability, and recovery
- Seeking: backend, search/retrieval, and systems internships
Mechanism-Conditioned Dynamics and Directional Effects in Iterative Retrieval
Status: Manuscript in preparation for a SIGIR full-paper submission; currently kept anonymous.
Question: When does ANN approximation become amplified—or damped—when retrieved evidence modifies future query states?
I developed a fidelity-conditioned coupled-trajectory framework that compares lower- and higher-fidelity retrieval processes while holding the corpus, encoder, initial query, feedback policy, retrieval depth, task, and evaluator fixed. The framework tracks how an ANN fidelity intervention propagates through query-state, candidate, and utility divergence across iterative feedback trajectories.
Key evidence:
- Across FEVER and HotpotQA, index-representation approximation expanded cross-fidelity separation, while IVF
nprobeand HNSWefSearchcould separate query states yet contract downstream candidate and utility disagreement. - On 3,316 FEVER-E5 validation queries, an approximately severity-matched representation-vs-
nprobecomparison produced a pairedH3absdifference of +0.033739, 95% CI [0.031838, 0.035642]. - A frozen 8.84M-passage MS MARCO replication evaluated 3,490 validation queries, 44 policies, and four feedback updates, yielding 1,535,600 trajectory rows and preserving the policy-weighted representation-vs-search-effort ordering.
- A score-channel audit replaced ANN-derived feedback weights with exact dot-product weights over the same selected candidates; the tested representation effect remained positive, narrowing the role of score geometry.
The study uses frozen protocols and splits, query-level or query-cluster bootstrap uncertainty, hash-linked artifacts, falsification tests, and explicit separation of primary and post-primary analyses.
Claim boundary: the evidence supports mechanism-conditioned behavior under the tested iterative retrieval settings—not a universal instability law, a new feedback operator, long-horizon agent behavior, or an exact-search comparison.
A public manuscript and artifact will be linked when anonymity constraints permit.
An evidence-led, event-driven commerce backend built with Java 25 and Spring Boot 4.1.0, with correctness defined around durable database state rather than process-level success.
Engineering highlights:
- Made the PostgreSQL commit the correctness boundary for order state, stock deduction, cart deletion, purchase snapshots, and outbound event intent.
- Implemented payment idempotency using command identity, SHA-256 request fingerprints, response snapshots, database uniqueness, and pessimistic locking around terminal-state races.
- Built a durable Transactional Outbox with
SKIP LOCKEDclaims, ownership leases, retry and terminal-failure handling, administrative replay, consumer deduplication, and persisted DLT governance. - Exercised the system under Kubernetes, CloudNativePG, Kafka, and Redis failure scenarios, with metrics, logs, traces, alerts, runbooks, supply-chain controls, and Terraform-managed OCI infrastructure.
Verified evidence:
- 172 passed / 0 failed tests with 92.55% instruction / 82.75% branch coverage in the latest retained clean-clone application verification.
- Three-replica Outbox drill: 90 durable
PUBLISHEDrows ↔ 90 valid unique Kafka sequences, with no missing or duplicate values in the tested workload. - CloudNativePG primary-loss drill: 51.543 s client-visible write RTO; 72 captured successful acknowledgements reconciled with 0 missing, and an independent restore verified 3,864 application rows.
- Kafka one-node-loss drill reconciled 3,733 captured acknowledgements with 0 missing; Redis Sentinel observed a replacement master at approximately 18 seconds.
- OCI acceptance ended with restricted SSH exposure, no direct public TCP/8080 access, a Terraform
No changesplan, and a successful controlled-reboot verification.
Boundary: this is a locally and development-cloud-validated engineering portfolio, not evidence of production-scale or multi-zone reliability.
- Distinguish configuration, executed verification, workload-scoped observation, and proposed architecture as different claim levels.
- Reconcile durable state after faults instead of inferring correctness from process recovery alone.
- Preserve negative and mixed results when they narrow the defensible conclusion.
- Freeze evaluation rules before held-out analysis and keep protocols, hashes, assumptions, and boundaries reviewable.
- Languages: Java, Python, SQL, Bash, Terraform/HCL
- Backend & data: Spring Boot, JPA/Hibernate, PostgreSQL, Kafka, Redis, Flyway, REST, Testcontainers
- Retrieval & evaluation: Faiss, IVF-PQ/SQ8, HNSW, dense pseudo-relevance feedback, nDCG, paired and query-cluster bootstrap
- Platform: Docker, Kubernetes/kind, GitHub Actions, OCI, Prometheus, Grafana, OpenTelemetry, Trivy, Cosign
- Email: sdgjklmv@gmail.com
- Location: Taichung, Taiwan
- GitHub: ravan-chuang


