DATA · AI AGENTS · CROSS-INDUSTRY

NapoliData

17+ years in data. Five industries, banking to healthcare.

Senior Data & AI Engineer. Eight years building cloud data platforms — lakehouse on Databricks and AWS, warehouse modeling in Snowflake and dbt, multi-agent systems on Bedrock — shipped inside regulated environments.

pipeline · live view
17+YEARS
8ROLES
5INDUSTRIES
WHERE I'VE SHIPPED: JOHNSON & JOHNSONBANCO GALICIAPRISMA MEDIOS DE PAGOBANCO PATAGONIAAPRENDE INSTITUTEFIVVYPROACTIVITI
// TRAJECTORY

The trajectory.

Eight roles — seven employers plus my own practice, in reverse order. Every entry maps to a LinkedIn role and a stack you can dig into on a call.

PROGRESS
0%
NOV 2025 — PRESENT
01 / 08

NapoliData LLC · Independent

Senior Data & AI Agent Engineer. Independent contractor providing data engineering and AI agent development services to clients in regulated industries. Work involves LLM-based observability, multi-agent orchestration, retrieval-augmented systems, and VPC-private model inference for environments with data-residency or compliance requirements.

AWSAzureDatabricksSnowflakeAnthropic ecosystem · MCPAirflow on K8sTerraformdbt

Engagements covered under NDA — architecture and outcomes discussable on a call.

JUN 2024 — NOV 2025
02 / 08

Proactiviti

Data Engineer. Cross-cloud ingestion from REST API and Kafka sources into Databricks, Delta Lake, Unity Catalog, Snowflake and Azure SQL — with data contracts enforced as CI/CD gates in GitLab on every pull request. Databricks Workflows and Airflow on Kubernetes with per-DAG images to cut cold-start overhead. Redshift cost down via sort/dist key redesign, materialized dbt models and query rewrites. Built a LangChain + Claude assistant that parsed pipeline logs into ranked root causes for on-call review.

−20%unplanned incidents · LLM monitoring agents
JUL 2023 — JUN 2024
03 / 08

Fivvy

AWS Data Engineer. Redshift analytical performance improved through dimensional-model redesign, materialized views and workload-management tuning. Pipelines converted into parameterized factories with GitLab CI/CD validation, so any team member could ship a change through review.

−30%S3 storage cost · lifecycle + partitioning
MAR 2022 — JUL 2023
04 / 08

Aprende Institute

AWS BI Data Engineer. Led the migration off SQL Server and SSIS onto AWS, rebuilding ETL on Glue, Lambda and Step Functions and serving analytics through Redshift, Athena and QuickSight. Cost cut by tuning Glue worker types, enabling job bookmarks for incremental loads and rescheduling production windows.

−35%processing cost
−50%manual reporting
+20%marketing ROI
JUL 2021 — MAR 2022
05 / 08

Johnson & Johnson · US Remote

Senior Data Engineer. Productionized data-science models: notebooks refactored into production Python, inference on Lambda + Glue, version control, unit testing and CloudWatch/EventBridge monitoring discipline. Large-scale ETL on AWS (EMR, Athena, S3, Redshift, Kinesis).

NOV 2020 — JUN 2021
06 / 08

Prisma Medios de Pago

Data Scientist Project Leader. Owned roadmap and delivery of Big Data & Analytics projects for Argentina's Visa acquirer across fraud, risk and merchant analytics. S3 · Athena · PySpark · Docker.

FEB 2016 — OCT 2020
07 / 08

Banco Galicia

Data Analyst → Senior Data Scientist (Marketing & Credit Risk). Propensity, cross-sell and segmentation models across insurance and lending; a lending qualification engine, a real-time recommender on Oracle, and an NLP ReMarketing chatbot. Measurable sales lift, material call-center savings.

Case study overview →
NOV 2009 — FEB 2016
08 / 08

Banco Patagonia

Data Analyst, Credit Risk Management. Seven years building and maintaining credit scoring models for retail and corporate portfolios — consumer loans, cards and SME credit lines. Python · SPSS · SQL.

// HOW I WORK

Three shapes of project.

Where the data comes from, what gets built, what ships.

01 · OBSERVABILITY · AI AGENTS01

From log noise to actionable incidents.

−20%unplanned incidents · Proactiviti
IN
Input

Airflow task logs, cloud alarms, dbt test failures, on-call Slack noise.

BL
Build

Multi-agent triage over normalized events: correlation, severity scoring, runbook retrieval via RAG — inference kept inside client infrastructure.

OUT
Output

Triaged notification with root-cause hint, linked runbook, next action. Fewer pings, faster MTTR.

02 · DATA ENGINEERING · CLOUD02

Heterogeneous sources into a queryable warehouse.

−30%S3 · −35% processing · Fivvy & Aprende
IN
Input

Postgres, S3 dumps, SaaS APIs, Kafka streams. Mixed schemas, mixed cadences, no shared dictionary.

BL
Build

Modular ETL/ELT on AWS or Azure, Spark on EMR/Databricks for heavy load, dbt for modeling and tests. Cost and partitioning tuned from day one.

OUT
Output

Curated tables in Redshift or Snowflake. BI-ready, SLA-tracked, documented. Storage and compute cut by a third.

03 · PREDICTIVE MODELING · MLOps03

From raw transactions to scored populations.

+20%marketing ROI · Aprende Institute · modeling practice built over 10 yrs in banking & insurance
IN
Input

Transactional warehouse, CRM attributes, behavioral signals, campaign history.

BL
Build

Feature engineering plus classical models (logistic, GBM, RFM, unsupervised), or embeddings and RAG when text dominates. MLflow tracking, batch inference in production.

OUT
Output

Scored customers piped to CRM, campaigns, credit decisions and churn watchlists. Decisions move from gut to evidence.

ILLUSTRATIVE REPLAY · TRIAGE AGENTidle
$ press Run triage to replay an incident through the agent graph
Synthetic walkthrough of the agent graph — the shipped version ran on client infrastructure.
// STACK

Stack I connect.

Not a laundry list. Tools shipped to production, grouped by where they sit in the data path. Highlighted = primary in the last two years.

CLOUD · PIPELINES

AWS GlueStep FunctionsAirflow on K8sLambdaEMRAthenaRedshiftAzure ADFSynapseKafkaKinesisCloudWatchEventBridgeIAMREST APIs

LAKEHOUSE · STREAMING

DatabricksPySparkDelta Live TablesUnity CatalogDelta LakeStructured StreamingAuto LoaderCDCDatabricks WorkflowsApache Iceberg

WAREHOUSE · ANALYTICS ENGINEERING

SnowflakedbtSQLPythonDimensional modelingData contractsSnowparkStar schemas · SCDIncremental · idempotencyPostgreSQLOracle

AI · LLMs · AGENTS

Anthropic ClaudeAmazon BedrockLangGraphMCPRAGLangChainOllamaHugging FacepgvectorSemantic cachingDynamic model routingLLMOpsNLPscikit-learnMLflow

PLATFORM · DELIVERY

TerraformKubernetesGitLab CI/CDDockerGitHub ActionsCloudFormationData qualityLineage · governance

BI · DELIVERY

QuickSightPower BITableauMetabase

CREDENTIALS

Google Cloud — MLOps
AWS Certified Cloud Practitioner
Postgraduate, Project Management — UTN FRBA
Accounting degree — the road into credit risk, then data science
Fluent English — client-facing on US teams since 2021 · Spanish native
// CONTACT

Let's talk.

A 30-minute call to see if there's a fit. No pitch.

PROJECT-BASED

Defined scope and timeline. Typically 4–12 weeks.

FRACTIONAL

Senior part-time capacity, embedded in your team. Typically 2–3 days a week.

ADVISORY

Architecture, code review, technical decisions. Monthly retainer — async plus a weekly call.

Buenos Aires (UTC−3) · 6+ hours overlap with US Eastern · Contracting through Napoli Data LLC · I reply within one business day

Third-party proof: read the recommendations on LinkedIn →