SENIOR BACKEND ENGINEER · NODE.JS / TYPESCRIPT · MYSQL

I build backends that scale before they have to.

Nearly seven years designing high-volume transaction systems — event-driven services, distributed transactions that stay consistent, and MySQL schemas that hold up when traffic triples — with PostgreSQL and MongoDB in the toolkit when the workload calls for them. This page is the part a resume can't carry: the architectures, the tradeoffs, and why I chose them.

vaishnavinarang4@gmail.comDownload resume ↓github/VN04 ↗linkedin ↗

Saharanpur, India (UTC+5:30) · open to remote & hybrid · senior / staff

6+ Yrs

Production backend engineering

10M+

Embeddings served at sub-100ms

15%

P99 latency cut on critical workflows

20%

Database load reduced under 3x growth

60%

Alert fatigue reduction (Velyn)

Experience

Oct 2025 — present

Tech Consultant

Clikd.ai

  • —Architected full-stack features for an AI recruitment platform, including document-processing pipelines that turn unstructured resumes and job descriptions into clean structured data — thousands of concurrent transformations with guaranteed consistency.
  • —Built semantic search on PostgreSQL/pgvector serving 10M+ embeddings at sub-100ms latency, which came down to index design and honest tradeoffs between indexing strategies.
  • —Owned backend deployment and operations on AWS (Lambda, RDS, SQS, EKS): JWT auth, queue-based processing for high-volume work, Sentry monitoring, reliability end to end.

Apr 2024 — Oct 2025

Lead Software Development Engineer

SaaS Labs

  • —Led the backend team on Node.js/NestJS services for a high-volume SaaS platform, scaling the architecture as the product went from 15 to 200+ concurrent users — establishing the patterns and standards that let the whole team scale, not just my own code.
  • —Diagnosed and resolved bottlenecks in the most critical workflows for a 15% P99 reduction: query profiling, serialization patterns, connection pooling.
  • —Architected event-driven workflows on Redis Pub/Sub that cut connection delays 40%, and added circuit breakers so external-service failures degraded gracefully instead of cascading.
  • —Built the reliability foundation — on-call rotation, documented incident response, post-incident reviews, SonarCloud gating in CI — and mentored junior engineers on API design and production ownership.

Apr 2022 — Mar 2024

Senior Software Development Engineer

SaaS Labs

  • —Led the migration of core services from PHP to Node.js/NestJS to support forecast 3x traffic growth, achieving 20% database load reduction through query optimization and pooling rather than infrastructure spend.
  • —Built analytics infrastructure on ClickHouse — schema design and complex SQL optimization for workloads that behave nothing like transactional ones.
  • —Architected distributed communication workflows that held data consistency under high volume, with asynchronous processing, failure recovery and graceful degradation.

Jan 2020 — Mar 2022

Software Development Engineer

SaaS Labs

  • —Built REST APIs and backend services for communication and sales-engagement products, including high-volume asynchronous workflows on Redis and Elasticsearch.
  • —Contributed to CI/CD infrastructure and NGINX-based deployments, managing staging environments and release processes.

Selected work

01Side project · full stack

Billeaf

Solo — architecture, backend, frontend

ProblemIndian shop owners, most without technical training, needed billing and inventory software they could actually operate. Existing tools were overengineered — the power was there, the usability wasn't.

OutcomeA SaaS where concurrent billing, stock movement and ledger writes never drift apart, behind an interface a shopkeeper learns in an afternoon.

NestJS · React · MySQL · Redis · saga orchestration · ACID transactions

Architecture

React SPA

shop terminal

↓

NestJS API

auth · validation

↓

Saga orchestrator

compensating steps

↓

Billing svc

invoices

Inventory svc

stock moves

Ledger svc

double entry

↓

MySQL

InnoDB · row locks

Redis

idempotency keys

Saga orchestrator owns the cross-service transaction; each service keeps its own consistency boundary.

02Side project · platform

Velyn

Solo — ingestion pipeline, clustering, API

ProblemProduction teams were drowning in alerts. Engineers spent days triaging errors that were variants of ones they'd already seen, so systemic causes never got fixed — the team just got faster at treating symptoms.

OutcomeAlert fatigue down 60%. Related errors collapse into one issue with a shared root cause, so the queue reflects problems rather than events.

Node.js · Python · npm + PyPI SDKs · MySQL · Redis · fingerprint clustering

Architecture

npm SDK

Node clients

Python SDK

Python clients

↓

Ingest API

accept · rate limit

↓

Redis queue

buffer · retries

↓

Cleaning / normalize

strip noise · parse traces

Fingerprinter

cluster · dedupe

↓

MySQL

issues · events

↓

Dashboard

alerts · trends

Ingestion is decoupled from processing by a queue; events are cleaned and normalized before anything downstream reads them.

03Consulting · Clikd.ai

Semantic search over 10M+ embeddings

Backend architecture, indexing, AWS ops

ProblemAn AI recruitment platform needed to match unstructured resumes against job descriptions semantically, at interactive latency, without standing up and paying for a separate vector database.

OutcomeSub-100ms query latency across 10M+ embeddings on PostgreSQL/pgvector, with thousands of concurrent document transformations landing consistently.

PostgreSQL + pgvector · HNSW · AWS Lambda · SQS · RDS · EKS · Sentry

Architecture

Upload API

resumes · JDs

↓

SQS

work queue

↓

Parser Lambda

unstructured → schema

Embedding worker

batch encode

↓

Postgres + pgvector

HNSW index

↓

Search API

JWT · filters

↓

Ranked matches

<100ms p99

Document processing is async and queue-driven; the read path stays a single indexed Postgres query.

04Production · SaaS Labs

PHP → Node/NestJS migration

Technical lead

ProblemThe core services were PHP and the product was forecast to take 3x traffic. Scaling infrastructure alone would have tracked cost linearly with growth, and the codebase couldn't support the async workflows the roadmap needed.

Outcome20% database load reduction through query and pooling work — capacity for 3x traffic without 3x infrastructure. 15% P99 latency cut on the critical workflows. Zero-downtime cutover.

NestJS · TypeScript · PostgreSQL · Redis Pub/Sub · ClickHouse · NGINX · circuit breakers

Architecture

Clients

web · API

↓

NGINX

strangler routing

↓

PHP monolith

shrinking surface

NestJS services

migrated domains

↓

Redis Pub/Sub

event fan-out

↓

PostgreSQL

pooled · indexed

ClickHouse

analytics

Strangler pattern: NGINX routes each migrated endpoint to Node while the monolith keeps serving the rest.

How I work

Most of the job isn't writing the service. It's deciding what the service is allowed to assume.

01

Correctness is a design decision, not a code review finding

Concurrency bugs and lost transactions come from ambiguous ownership of state. I decide where each invariant lives before writing the service, so the code has one obvious place to enforce it.

02

Measure before optimizing, then measure the thing users feel

A 15% P99 cut came from profiling real queries, not from guessing. Averages hide the requests that make people distrust your product; I tune the tail.

03

Design for the failure, not the happy path

External services will time out. Retries will duplicate. I assume both — idempotent processing, circuit breakers, compensating steps — so degradation is a feature rather than an incident.

04

The architecture that only I understand is a liability

I write the decision down with its cost, so the next engineer can tell whether the constraint still holds. Most of what I documented above started as a note for my own team.

Leading a team

I took a backend team through a product going from 15 to 200+ concurrent users. The scaling work that mattered wasn't mine alone — it was establishing the architectural patterns and review standards that let the rest of the team make the same call I would have, without asking me.

Set up the on-call rotation and incident-response runbook from nothing, ran blameless post-incident reviews, and wired SonarCloud into CI so quality arguments happened at the pipeline instead of in review comments.

15 → 200+

Concurrent users the architecture I led grew to support

Mentoring

Junior engineers, on API design and production ownership

On-call

Rotation, incident runbook and post-incident review process, established from scratch

Stack

Languages

  • TypeScript
  • Node.js
  • JavaScript
  • Python
  • SQL

Frameworks

  • NestJS
  • Express.js
  • React
  • REST
  • GraphQL

Data

  • MySQL
  • PostgreSQL
  • pgvector
  • ClickHouse
  • MongoDB
  • Redis
  • Elasticsearch

Cloud & infra

  • AWS Lambda
  • RDS
  • SQS
  • EC2
  • EKS
  • Docker
  • Kubernetes
  • CI/CD
  • IaC

Architecture

  • Microservices
  • Event-driven
  • Saga / distributed txn
  • Queue-based systems
  • Idempotency
  • Circuit breakers
  • Connection pooling

Practice

  • SOLID
  • OOP
  • Automated testing
  • Code review
  • SonarCloud
  • Load testing
  • Sentry
  • New Relic
  • Incident response

Integrated B.Tech + M.Tech, Computer Science — Gautam Buddha University, 2015–2020 · CGPA 8.99/10

If you're hiring for systems that have to stay up, let's talk.

Open to senior and staff backend or full-stack roles — remote or hybrid. Happy to walk through any of the architectures above in detail.

vaishnavinarang4@gmail.com