Software Development Engineer · Open for Full Time Roles

Deevna
Reddy

Building systems that hold up under real world constraints.
Specializing in Backend Systems.

Alexa Calendar50M+Users Supported
DSAR Compliance70%Manual Ops Reduced
JDK 17 Upgrades15+Zero Downtime Rollouts
Research Citations2x PapersSpringer & IEEE Published
01.

About

Hi, I'm Deevna! I build backend systems and ML driven products that hold up under real world constraints, whether that's compliance deadlines, audit cycles, or millions of users hitting a service at once.

I studied Computer Science Engineering (AI & Data Analytics) at Sri Ramachandra Institute of Higher Education & Research in Chennai. Most of what I know beyond that came from shipping things at Amazon: closing out compliance risk on a service used by 50M+ people, migrating a stack of microservices with zero downtime, teaching an on device model to read handwriting off a Kindle Scribe.

Outside of work, I like building agents and systems that make messy, unstructured problems tractable: claims documents, fraud patterns, and clinical healthcare datasets. My research has been published in IEEE ICCDS 2025 (First Author) and Springer Nature.

QUICK FACTS

📍Chennai, India
🎓CS Engineering, AI & Data Analytics
💼SDE @ Amazon
🧠Backend · ML / RAG
📄2x Published (Springer & IEEE)

Technical Skills & Architecture Matrix

Classified Pillars
Core Languages

Java, Python, C++, TypeScript, SQL, Bash

Data & Applied ML

PyTorch, Scikit learn, Multi Agent RAG, SHAP, LiDAR Point Cloud Processing, FastEmbed, NLP

Cloud & Distributed Scale

AWS (Lambda, S3, DynamoDB, EC2), Docker, CI/CD, Policy Engine, Linux, Microservices

02.

Experience

Software Development Engineer

Jan 2025 to Mar 2026

Amazon · Chennai, IN 🇮🇳

Started as an intern building Python ETL pipelines that turned raw Kindle Scribe stroke data into a per word feature dataset for handwriting recognition training, then ported the stroke straightening model to native C++ for real time, dependency free on device inference, improving baseline detection robustness by 25%. Owned DSAR compliance onboarding for Alexa Calendar end to end, cutting manual data access and deletion processing time by 70% for a service supporting 50M+ users. Closed every outstanding policy engine risk item as sole contributor across two audit cycles, and led 15+ JDK 8/11 → 17 migrations with zero downtime rollouts, improving runtime performance by 20%.

  • Java
  • AWS
  • JDK 17
  • CI/CD
  • Policy Engine
  • Python
  • C++

AI & ML Intern

Aug 2023 to Oct 2023

Agilisium Consulting · Chennai, IN 🇮🇳

Built an AI powered HR chatbot using NLP and RAG over internal knowledge sources, cutting manual HR query workload by 40%.

  • Python
  • NLP
  • RAG

Data & Image Analyst

Sep 2023 to Apr 2024

REUDE Technologies · Chennai, IN 🇮🇳

Worked on drone based aerial data acquisition and point cloud processing, improving spatial data accuracy by 20% over raw sensor output.

  • Drone Data
  • Point Cloud Processing
03.

Projects

Problem, approach, outcome, and what I'd change. Numbers only where I can explain how they were measured.

AI/ML · MULTI AGENT01

Claims Agent

Turning messy claim documents into structured, auto routed decisions.

15+

FIELDS EXTRACTED

80%

AUTO TRIAGED

100%

PII MASKED

1

HUMAN REVIEW GATE

FIELD SCHEMA · 9 / 9 EXTRACTED

Policy #

Claimant

Incident date

Damage type

Est. cost

Coverage

Adjuster

Priority

Status

PIPELINE STAGES

Status: Ready
1. Ingest Raw PDF
Pending
2. Extract 15 Structured Fields
Pending
3. LLM Validation Gate
Pending
4. Auto Triage Queue
Pending
“The 20% that do not auto route are not random, they cluster on ambiguous damage type fields, which is exactly where a human should be looking anyway.”

PROBLEM

Claims teams drown in unstructured PDFs, scans, and emails containing sensitive policyholder PII. Every claim needed a human just to redact records and figure out what it said before anyone could decide anything.

APPROACH

An agent pipeline (Ingest & Mask PII → Extract → Validate → Route → Queue) sanitizes sensitive policyholder records, pulls 15+ structured fields out of raw documents, then runs an LLM driven validation pass before routing, so low confidence extractions never slip through silently.

OUTCOME

80% of first pass triage now runs without a human touching it, with full PII compliance. The other 20% get flagged with the specific field that failed validation, not just a generic 'needs review' tag.

WHAT I'D CHANGE

The validation step catches malformed fields, not ambiguous ones: a claim that is technically well formed but contextually wrong still passes. I would add a second cross field consistency check before scaling past pilot volume.

PythonLLMsRAGMulti AgentPII RedactionData GovernanceView source
FINTECH · FRAUD DETECTION02

TradeShield

Catching fraud patterns a logistic regression would miss.

+25%

PRECISION RECALL AUC

<1s

SCORING LATENCY

4

RISK TIERS

24/7

LIVE SCORING

PRECISION RECALL AUC

Baseline

TradeShield

+25% AUC

RISK TIERS · LIVE SCORING

Low Risk (0.04)
low riskreviewhigh riskauto block
“The behavioral model scores sequences, not single transactions: the same account making five small, spaced out transfers is where the baseline went blind.”

PROBLEM

A plain logistic regression baseline flagged obvious fraud fine, but missed behavioral patterns: the same account making small, spaced out, plausible looking transactions that only look wrong in aggregate.

APPROACH

Pattern recognition on transaction features, layered with a behavioral anomaly model that scores sequences rather than single transactions. Served via FastAPI so a flagged transaction shows up on the risk dashboard within a second, not the next morning's batch job.

OUTCOME

Precision recall AUC improved 25% over the logistic regression baseline, with the biggest gains on exactly the slow burn fraud patterns the baseline missed.

WHAT I'D CHANGE

The behavioral model needs a longer transaction history to be confident, since new accounts get scored on thinner evidence. I would add a separate cold start tier instead of forcing new accounts through the same model.

PythonScikit learnFastAPINext.jsView source
PUBLISHED · IEEE ICCDS 202503

Genomic & Imaging Biomarker Analysis

Published at IEEE ICCDS 2025.

98.32%

FUSION ACCURACY

0.99

AUC

3

CLASSES: LUAD/LUSC/NORMAL

1st

AUTHOR CREDIT

TOP SHAP FEATURES

EGFR94%
TP6381%
NKX2.173%
Grad CAM nodule region68%

MODALITIES FUSED

DenseNet121 CT branchRF + SHAP gene branch512D fused
“EGFR signaling topped the pathway enrichment, matching known lung cancer biology, which is what actually made the accuracy number trustworthy.”

PROBLEM

Lung cancer AI tools typically use one data type and give a bare label. Clinicians cannot act on a black box call, especially on small, ambiguous nodules where a biopsy is invasive and a miss is costly.

APPROACH

A dual branch model: DenseNet121 + Grad CAM reads CT scans (LIDC IDRI, ~53K scans) for visual explanations, while a Random Forest + SHAP branch reads TCGA RNA seq gene expression (~798 samples) for molecular explanations. Late fusion combines both into one classifier predicting LUAD, LUSC, or Normal.

OUTCOME

The fused model hit 98.32% accuracy and 0.99 AUC, beating either modality alone (90.15% imaging only, 95.47% genomics only). SHAP flagged EGFR, TP63, and NKX2.1 as top drivers, consistent with known cancer pathways. First author paper, IEEE ICCDS 2025.

WHAT I'D CHANGE

Validation was retrospective on public datasets (LIDC IDRI, TCGA), not real clinical cases. I would want prospective validation with a hospital partner before trusting this near an actual diagnosis.

DenseNet121Grad CAMRandom ForestSHAPView paper
PUBLISHED · SPRINGER NATURE04

Ethnic Disparities in ASD Analysis

Statistical analysis and machine learning evaluation of ethnic disparities in Autism Spectrum Disorder among toddlers.

Springer

NATURE CHAPTER

Toddlers

COHORT FOCUS

ASD

CLINICAL SCREENING

Author

PUBLICATION CREDIT

TOP SHAP FEATURES

EGFR94%
TP6381%
NKX2.173%
Grad CAM nodule region68%

MODALITIES FUSED

DenseNet121 CT branchRF + SHAP gene branch512D fused
“EGFR signaling topped the pathway enrichment, matching known lung cancer biology, which is what actually made the accuracy number trustworthy.”

PROBLEM

Traditional ASD screening models often underrepresent minority ethnic cohorts, leading to delayed diagnoses and disparate screening outcomes during early childhood development.

APPROACH

Evaluated demographic and behavioral screening indicators across ethnic cohorts using statistical hypothesis testing and supervised classification to measure disparate impact.

OUTCOME

Quantified significant variance in screening markers across ethnic groups, establishing benchmarks for equitable pediatric diagnostic frameworks. Published in Springer Nature.

WHAT I'D CHANGE

I would expand the longitudinal cohort data to track post intervention developmental outcomes across multiple clinical hospital networks over a multi year timeframe.

Statistical ModelingMachine LearningHealthcare DataSpringer NatureView paper
04.

Contact

Say hello!

I'm always open to new challenges and collaborations. Whether you have a question or just want to say hi, I'll get back to you!

PHONE

+91 9940266618

SOCIAL

© 2026 Deevna Reddy