# Adil Munawar — full profile > Complete machine-readable profile of Adil Munawar: machine-learning engineer for agricultural remote sensing and full-stack developer in Lahore, Pakistan. Generated from the data behind https://adilmunawar.vercel.app; the concise version is https://adilmunawar.vercel.app/llms.txt and the JSON version is https://adilmunawar.vercel.app/ai-profile.json. ## Profile - Name: Adil Munawar - Also known as: AdilMunawarX, Adil Khokhar - Headline: Machine-learning engineer for agricultural remote sensing · full-stack developer - Summary: I train segmentation and time-series models on satellite imagery to map fields and crops, and build the web products that put those maps in front of people. - Roles: Machine Learning Engineer — agricultural remote sensing; Project Lead & Web Developer at Nexsus Orbits; Database Administrator at AOS - Organisations: Nexsus Orbits (https://nexsusorbits.com); AOS - Education: Govt. Islamia Graduate College Civil Lines, Lahore - Location: Lahore, Punjab, Pakistan (UTC+5) - Work mode: remote worldwide - Availability: Available for new engagements; remote worldwide. Email is best; replies within a day. - Canonical URL: https://adilmunawar.vercel.app - Profiles: https://github.com/AdilMunawar, https://www.linkedin.com/in/adilmunawar/, https://x.com/adilmunawarx, https://dev.to/adilmunawar, https://leetcode.com/u/AdilMunawar/, https://www.instagram.com/adilmunawarx/ ## Specialities - Agricultural remote sensing on Sentinel-2 and high-resolution imagery (NDVI/EVI phenology, cloud masking). - Crop-type classification from satellite time series (1D-CNN, LSTM, temporal attention). - Field boundary delineation with an HRNet-W48 segmentation model for Zaraat Dost Private Limited (Kishtwar region), producing GIS-ready parcels. - Farm digitization pipelines: imagery to clean vector parcels in PostGIS, with human-in-the-loop QA. - Crop yield prediction with XGBoost and SHAP explanations. - Enterprise RAG pipelines (pgvector, hybrid retrieval, rerankers, cited answers). - Agentic systems and custom Model Context Protocol (MCP) tool servers with scoped permissions and audit logs. - Full-stack products on Next.js/React + Supabase, deployed on Vercel. ## Services Four kinds of work taken on, from a single model to a shipped product. ### 01 Machine learning & remote sensing You get: null ### 02 RAG & agentic systems You get: null ### 03 Full-stack products You get: null ### 04 Data & cloud engineering You get: null ## Projects 24 projects. Entries marked "model" are private client or internal R&D work: they are described here but have no public repository or live URL, and no metrics or dates are claimed for them. Products and repositories list their public links. ### Agri-Tech & Geospatial #### HRNet-W48 Field Boundary Delineation, Kishtwar - Kind: model (highlighted on the site) - Domain: Agri-Tech - Client: Zaraat Dost Private Limited - Architecture: HRNet-W48; task: Field Boundary Segmentation; framework: PyTorch - Tech: PyTorch, HRNet-W48, Rasterio, GDAL, GeoPandas, QGIS - Links: none (built for Zaraat Dost Private Limited, no public repository) Semantic segmentation of agricultural parcel boundaries from high-resolution satellite imagery of the Kishtwar region, using an HRNet-W48 backbone that keeps full-resolution feature maps so thin field edges survive downsampling. Trained with a boundary-aware loss on digitised reference parcels, with tiled and overlap-stitched inference, morphological cleanup and vectorisation of the boundary raster into GIS-ready polygons. #### Crop-Type Classification from Satellite Time Series - Kind: model (highlighted on the site) - Domain: Agri-Tech - Client: Private client - Architecture: Temporal CNN / LSTM; task: Crop-Type Mapping; framework: PyTorch - Tech: Sentinel-2, PyTorch, 1D-CNN / LSTM, Temporal Attention, Google Earth Engine, xarray - Links: none (private client work, no public repository) Per-parcel crop-type mapping from Sentinel-2 time series, built on cloud-masked NDVI and EVI phenology stacks resampled to a regular temporal grid across the growing season. Compares 1D-CNN, LSTM and temporal-attention classifiers on the sequences, then aggregates pixel predictions to parcel level to produce a crop map with per-class confidence. #### Farm Digitization & Parcel Mapping Pipeline - Kind: model - Domain: Geospatial - Client: Private client - Architecture: Imagery → Vector Pipeline; task: Farm Digitization; framework: Python + PostGIS - Tech: PostGIS, GeoPandas, Shapely, GDAL, QGIS, Python - Links: none (private client work, no public repository) End-to-end pipeline that turns raw satellite imagery into clean vector farm parcels: preprocessing and mosaicking, model-based boundary extraction, polygonisation, topology fixes and attribute joins. Includes a human-in-the-loop QA workflow for reviewing and correcting parcels, with outputs delivered as GeoJSON and Shapefiles and stored in PostGIS for downstream analytics. #### NDVI Change Detection & Field Monitoring - Kind: model - Domain: Agri-Tech - Client: Private client - Architecture: Index Time Series + Thresholding; task: Change Detection; framework: Python - Tech: Sentinel-2, Google Earth Engine, Rasterio, Pandas, FastAPI, PostGIS - Links: none (private client work, no public repository) Multi-date vegetation-index monitoring over digitised parcels, computing NDVI and related indices per field on each clear Sentinel-2 acquisition and flagging significant deviations from the field's seasonal baseline. Change events and time-series charts are exposed through a lightweight dashboard so agronomists can prioritise field visits. #### Cloud & Shadow Masking for Sentinel-2 - Kind: model - Domain: Geospatial - Client: Internal R&D - Architecture: U-Net; task: Cloud / Shadow Segmentation; framework: PyTorch - Tech: PyTorch, U-Net, Sentinel-2, Rasterio, NumPy - Links: none (internal R&D, no public repository) Pixel-wise cloud and cloud-shadow segmentation for Sentinel-2 scenes, trained on multispectral bands to produce cleaner masks than the standard scene-classification layer for agricultural time-series work. Used as a preprocessing step so that only valid observations enter the phenology stacks that feed the crop classification models. #### Satellite Imagery Deep Learning - Kind: repo - Domain: Geospatial - Architecture: U-Net / Mask R-CNN / YOLO; task: Earth Observation; framework: PyTorch - Tech: Deep Learning, Segmentation, Object Detection, Remote Sensing - Links: code: https://github.com/Adilmunawar/Satelite-Imagery-Deep-Learning Curated reference of deep learning architectures for Earth observation: classification, semantic and instance segmentation, and object detection on satellite and aerial imagery, organised by task and domain. ### Machine Learning #### Crop Yield Prediction with XGBoost - Kind: model - Domain: Agri-Tech - Client: Private client - Architecture: Gradient Boosted Trees; task: Yield Regression; framework: XGBoost - Tech: XGBoost, scikit-learn, SHAP, Optuna, Pandas - Links: none (private client work, no public repository) Gradient-boosted regression model for field-level yield estimation from tabular agronomic records combined with remote-sensing features such as seasonal NDVI statistics and phenology dates. Uses grouped cross-validation by season and region to avoid leakage, with SHAP values used to explain which features drive each prediction. #### CNN Land-Cover & Crop Classification - Kind: model - Domain: Geospatial - Client: Internal R&D - Architecture: ResNet (multispectral); task: Land-Cover Classification; framework: PyTorch - Tech: PyTorch, ResNet, Albumentations, Focal Loss, ONNX - Links: none (internal R&D, no public repository) Patch-based convolutional classifier for land-cover and crop classes on multispectral satellite tiles, fine-tuned from an ImageNet-pretrained ResNet with the first convolution adapted to additional bands. Handles severe class imbalance with weighted sampling and focal loss, and reports per-class metrics rather than a single accuracy figure. #### LSTM Sequence Models for Phenology & Anomaly Detection - Kind: model - Domain: Agri-Tech - Client: Private client - Architecture: Bi-LSTM Autoencoder; task: Phenology & Anomaly Detection; framework: TensorFlow - Tech: TensorFlow, Keras, LSTM / GRU, NumPy, MLflow - Links: none (private client work, no public repository) Recurrent models over per-field vegetation-index sequences that learn the expected seasonal curve and identify key phenological stages such as emergence, peak and senescence. A reconstruction-error variant flags irrigation gaps and other anomalies when a field's trajectory departs from its learned pattern. #### Model Serving & Scheduled Inference - Kind: model - Domain: Geospatial - Client: Private client - Architecture: Containerised Inference; task: MLOps Deployment; framework: FastAPI + Docker - Tech: Docker, FastAPI, ONNX Runtime, PostgreSQL, Cron / Airflow, MLflow - Links: none (private client work, no public repository) Deployment layer for the remote-sensing models: containerised FastAPI services for on-demand inference and scheduled jobs that pull new imagery, run segmentation or classification over areas of interest and write results to PostGIS. Includes model versioning, run logging and basic drift checks on input statistics. #### Jarvis Voice AI Assistant - Kind: repo - Domain: Product - Architecture: STT → Router → Skills; task: Voice Agent; framework: Python - Tech: Python, Speech Recognition, OpenCV, Selenium - Links: code: https://github.com/Adilmunawar/Jarvis Modular Python assistant built around speech-to-text and text-to-speech, with a central router that dispatches voice commands to automation, web, scheduling and computer-vision modules. ### LLM & Agents #### Enterprise RAG Pipeline - Kind: model (highlighted on the site) - Domain: LLM Systems - Client: Private client - Architecture: Hybrid Retrieval + Reranker; task: Grounded Q&A; framework: LangChain - Tech: pgvector, LangChain, Claude / OpenAI API, FastAPI, Redis, PostgreSQL - Links: none (private client work, no public repository) Retrieval-augmented question answering over a private company knowledge base: document parsing and structure-aware chunking, embedding ingestion into pgvector, and hybrid dense plus keyword search with a cross-encoder reranker. Answers are grounded with inline citations to source passages, and an evaluation harness tracks retrieval recall and answer faithfulness across releases. #### Agentic Workflows & MCP Tool Server - Kind: model - Domain: LLM Systems - Client: Private client - Architecture: MCP Tool Server + Agents; task: Agentic Automation; framework: TypeScript - Tech: TypeScript, MCP SDK, Zod, Claude API, PostgreSQL, Docker - Links: none (private client work, no public repository) Multi-step LLM agents that operate over internal APIs and databases through a Model Context Protocol server exposing typed, schema-validated tools and resources. Includes scoped permissions per tool, confirmation gates for side-effecting actions, and structured audit logs of every tool call for review. #### AdiGaze Resume Intelligence - Kind: product - Domain: Product - Tech: React, TypeScript, Supabase, AI Parsing - Links: live: https://adigaze.vercel.app; code: https://github.com/Adilmunawar/AdiGaze AI recruitment platform that parses unstructured resumes into structured candidate records and ranks applicants against job requirements, with MFA, external submissions and processing telemetry on Supabase edge functions. #### AdiGon AI Assistant - Kind: product - Domain: Product - Tech: TypeScript, React, Gemini, Supabase - Links: live: https://adigon.vercel.app; code: https://github.com/Adilmunawar/AdiGon-AI Conversational assistant that understands multi-format inputs (PDF, code, documents), with built-in voice recognition, a developer mode and Markdown rendering on a Supabase backend. #### AdiHunt SEO Content Intelligence - Kind: product - Domain: Product - Architecture: LLM Content Pipeline; task: SEO Generation; framework: Gemini - Tech: React, TypeScript, Gemini, Supabase - Links: live: https://adihunt.vercel.app; code: https://github.com/Adilmunawar/AdiHunt AI SEO platform generating long-form optimised articles with keyword research, competitor analysis and technical audits, plus content planning and team collaboration. #### AdiFy AI Resume Builder - Kind: product - Domain: Product - Tech: Next.js, Genkit, Gemini, Firebase - Links: live: https://adifyai.vercel.app; code: https://github.com/Adilmunawar/AdiFy AI-assisted resume builder with content refinement and professional templates, powered by Gemini through Genkit, including image cropping and export tooling. ### Products #### AdiFlux Image Generator - Kind: product - Domain: Product - Tech: Next.js, Genkit, Gemini, Firebase - Links: live: https://adiflux.vercel.app; code: https://github.com/Adilmunawar/AdiFlux Text-to-image generation web app with no login required, built on Google's Genkit framework with a streamlined component-based Next.js interface. #### AdiMage AI Photo Editor - Kind: product - Domain: Product - Architecture: Generative Editing; task: Image Enhancement; framework: GenAI SDK - Tech: React, TypeScript, Google GenAI - Links: live: https://adimage.vercel.app; code: https://github.com/Adilmunawar/AdiMage AI photo editing tool for object removal, background replacement and detail upscaling with template-based workflows, built on Google's GenAI SDK. #### AdiCorp HRMS Platform - Kind: product - Domain: Product - Tech: React, TypeScript, Supabase, jsPDF - Links: live: https://adicorp.vercel.app; code: https://github.com/Adilmunawar/AdiCorp Modular human resource management system with admin dashboard and employee self-service, covering attendance, leave, overtime, payroll and events, with PDF and spreadsheet export. #### AdiNox OTP Authenticator - Kind: product - Domain: Product - Tech: React, TypeScript, Supabase, face-api.js - Links: live: https://adinox.vercel.app; code: https://github.com/Adilmunawar/AdiNox Two-factor authentication app generating TOTP tokens with local persistence, QR scanning and face-api.js biometric access control. #### Aditron Social Chat - Kind: product - Domain: Product - Tech: React, TypeScript, Supabase, Realtime - Links: live: https://aditron.vercel.app; code: https://github.com/Adilmunawar/aditrondev Real-time social chat application with emoji support and OTP/QR utilities on a Supabase realtime backend. #### NUREH E-Commerce Storefront - Kind: product - Domain: Product - Architecture: Static Storefront; task: E-Commerce; framework: React + Vite - Tech: React, Vite, Vanilla CSS - Links: live: https://nureh.pk; code: https://github.com/Adilmunawar/Nureh Clothing storefront optimised for fast loads with LCP asset preloading and layout-shift prevention, featuring video reels, a cart drawer and product carousels. #### Federals.live News CMS - Kind: product - Domain: Product - Architecture: Headless CMS; task: Publishing; framework: React + Supabase - Tech: React, Vite, Supabase, PostgreSQL - Links: live: https://federals-live.vercel.app; code: https://github.com/Adilmunawar/Federals.live Production news CMS with a Markdown editor, SEO tooling, breaking-news ticker and a role-secured admin portal backed by Supabase row-level security. ## Toolkit - null: PyTorch, HRNet / U-Net, Temporal CNN & LSTM, scikit-learn, XGBoost, TensorFlow / Keras, ONNX Runtime, MLflow, Google Earth Engine, xarray - null: Python, Node.js, FastAPI, PostgreSQL / PostGIS, Supabase, Redis, pgvector, GDAL / Rasterio, GeoPandas / Shapely, NumPy / pandas - null: TypeScript, React, Next.js, Vite, Tailwind CSS, HTML & CSS, Zod, face-api.js, jsPDF, Framer Motion - null: Docker, AWS, Azure, Vercel, GitHub Actions, Cron / Airflow, RAG pipelines, MCP servers, LangChain, QGIS ## Certifications Courses and assessments from Google Cloud, AWS, Microsoft, Anthropic and LinkedIn. - Operationalising EU Space for Security — EUSPA (issued Aug 2026). EU Agency for the Space Programme course on operationalising EU space services for border security and geospatial intelligence. Credential ID 6966612178c62e5fe0048a0f. - MATLAB Onramp — MathWorks (issued Jul 2026). Hands-on introduction to MATLAB for numerical computing and data analysis. - Advanced CloudFormation: Macros — AWS (issued Jul 2026). Extending AWS CloudFormation templates with macros for reusable infrastructure as code. - Machine Learning with Python — MIT Professional Education (issued Jul 2026). Supervised and unsupervised learning foundations implemented in Python. Credential ID 4b7ac91e0f3d82c56a19fe0332d8a17. - Computational Probability and Inference — MIT Professional Education (issued Jun 2026). Probabilistic modelling and inference methods for data-driven systems. Credential ID e4ecfea7a7014b2483579dd1e7356c23. - Web Analytics by Accenture — Accenture / FutureLearn (issued Apr 2026). Web analytics for data-driven decisions, delivered by Accenture on FutureLearn. Credential ID v6zgddz. - Certified Software Engineer — HackerRank (issued Apr 2026). Verified software engineering and problem-solving proficiency. Credential ID b6411a6e46da. - Agile Foundations — LinkedIn / PMI (issued Apr 2026). Agile foundations in collaboration with the Project Management Institute. - SQL and Relational Databases 101 — Cognitive Class / IBM Skills Network (issued Mar 2026). Relational database design and SQL querying fundamentals. Credential ID 6acdbc63fde448e3a5d2304cd0333ea8. - Claude in Amazon Bedrock — Anthropic (issued Mar 2026). Integrating and optimizing Claude models on Amazon Bedrock, including the Model Context Protocol. Credential ID wfq8d7zjnka8. - Chronicle SOAR Developer — Google Cloud Skills Boost (issued Feb 2026). Building security orchestration, automation and response playbooks on Google Chronicle SOAR. Credential ID 22073529. - Building Language Models on AWS — AWS (issued Dec 2025). Training and deploying language models on Amazon Web Services. - Google Ads Apps Certification — Google Cloud Skills Boost (issued Dec 2025). Building and integrating applications with Google Ads APIs. Credential ID 169263387. - Azure Cloud Computing — Microsoft (issued May 2025). Architecting secure, scalable solutions on Microsoft Azure. Credential ID AdilMunawar4765. - Application Modernization with Google Cloud — Google (issued Mar 2025). Modernizing legacy architectures for performance, scalability and security. Credential ID 14164265. - LinkedIn Content and Creative Design — LinkedIn (issued Mar 2025). Technical content and creative design fundamentals. Credential ID zn7dbp7a2cw3. - MLOps with Vertex AI — Google (issued Feb 2025). Managing machine learning models at scale on Vertex AI. Credential ID 14116643. - Machine Learning Operations for Generative AI — Google (issued Feb 2025). MLOps workflows across the machine learning lifecycle for generative AI. Credential ID 14101465. - Advanced Webhook Concepts — Google Cloud. Advanced webhook concepts for real-time data synchronization between applications. - Advanced Performance Measurements — Google. Enterprise performance measurement for high-speed applications. - CCAI Frontend Integrations — Google Cloud. Contact Center AI frontend integrations with user-centric interfaces. ## Badges and learning paths - GitHub: GitHub Actions, GitHub Admin, GitHub Advanced Security, GitHub Agentic AI Developer - LeetCode: Top SQL 50, Algorithm Deconstructor, Architecture Builder, Data Navigator, Mathematical Insight, Introduction to Pandas, 100 Days 2025, 100 Days 2026, 50 Days 2025, 50 Days 2026 - Microsoft & AWS: Azure Virtual Machines, Azure Compute Resources, Lakehouse with Microsoft Fabric, MD-102 Endpoint Management, Entra Identity Protection, Azure Machine Learning Models, Azure Core Data Concepts, Azure Backup, DAX Time Intelligence, Microsoft 365 Defender, Azure Storage Security, AWS Generative AI Developer - Google: Android Studio User, Firebase Studio Community, Gemini CLI User, Google Cloud & NVIDIA Community, Google Cloud Innovator, Google Developer Program Premium, Google Generative AI Leader, DOM Detective ## Recognition - LeetCode Top 5% Global Solver - GitHub Achievement: Pull Shark (Gold) - GitHub Achievement: Pair Extraordinaire (Gold) - GitHub Achievement: Galaxy Brain (Silver) - Google Ads Apps Certification (2025) - AWS Building Language Models (2025) - Microsoft Azure Cloud Computing (2025) ## Activity Public commits and problem-solving over the last year, refreshed automatically by GitHub Actions. - LeetCode: 1335 problems solved (315 easy, 577 medium, 443 hard), acceptance rate 89.5%, global ranking 12,631 — https://leetcode.com/u/AdilMunawar/ - GitHub: 8,929 public contributions in the last year — https://github.com/AdilMunawar ## Case studies Long-form write-ups published at https://adilmunawar.vercel.app/#case-studies. Diagrams are given as Mermaid source. ### Delineating Terraced Field Boundaries in Kishtwar with HRNet-W48 - Stack: PyTorch, HRNet-W48, Rasterio, GeoPandas, QGIS - Summary: How we turned high resolution imagery of terraced farmland into a topology-clean parcel layer for Zaraat Dost, using an HRNet-W48 boundary model, overlap-stitched tiled inference, watershed instancing and shared-arc vectorisation. - Challenge: Kishtwar's terraced parcels are small, irregular and packed so tightly that neighbouring fields share crop, stage and texture; the only evidence of a boundary is a bund or riser one or two pixels wide. Zaraat Dost needed those boundaries as closed, non-overlapping polygons their GIS team could edit, not a probability raster, and standard encoder-decoder segmentation blurred the thin edges into merged strips. - Solution: We trained an HRNet-W48 model whose full resolution stream never drops the detail, with a boundary channel, an interior mask and a distance-weighted loss that concentrates the penalty at edges. Inference runs over overlapping rasterio windows blended with cosine ramp weights so no tile seam survives, followed by hysteresis thresholding, skeletonisation, marker-based watershed, and a vectorisation step that simplifies shared arcs once before rebuilding faces, delivered as GeoJSON and GeoPackage with per-parcel confidence. #### Delineating terraced field boundaries in Kishtwar with HRNet-W48 Zaraat Dost Private Limited needed parcel boundaries for farmland in the Kishtwar region, where cultivation climbs the valley sides in terraces. What they wanted was not a raster but a vector layer their GIS team could open, edit and attribute: one polygon per parcel, no overlaps, no gaps, closed rings. This is how we got from high resolution imagery to that layer, what we tried on the way, and what broke. ##### What made Kishtwar hard Terraced agriculture is close to the worst case for boundary extraction. Parcels are small, so a few pixels of error at the edge is a large fraction of the parcel, and irregular, following contour lines rather than cadastral grids. The boundaries are physically thin: a stone riser or grass bund that shows as a one or two pixel shadow line when the sun is low and disappears when it is high. Neighbouring terraces often carry the same crop at the same stage, so their interiors look identical; the only evidence that two parcels are two parcels is the line between them, and any method that segments interiors will merge neighbours. We had to model the boundary itself. ##### Why HRNet-W48 The usual segmentation backbones all throw away resolution somewhere in the middle. A U-Net compresses the image through its encoder and rebuilds detail with skip connections; DeepLabv3+ runs at an output stride of eight or sixteen and upsamples bilinearly. For a boundary one or two pixels wide, the decision about whether a pixel is a bund has already been pooled with its neighbours before the decoder sees it, and the decoder can only smear it back out. Our early U-Net boundary maps looked plausible at a glance and were wrong at the terraces: risers between narrow strips fused into one blurred edge, and the watershed downstream merged the strips. HRNet keeps a high resolution stream alive from the first stage to the last and adds progressively lower resolution streams in parallel, at one eighth, one sixteenth and one thirty-second of the input. At the end of every stage the streams exchange information: each receives strided convolutions of the finer streams and upsampled maps of the coarser ones. The fine stream never has to reconstruct detail; it always had it, and it gains context from the coarse streams without paying in resolution. The W48 variant widens every stream's channels relative to W18 and W32. It costs memory per tile, but inference here is an offline batch job, not a real time service, so we spent the budget on capacity rather than latency. ```figure hrnet-kishtwar-boundaries/streams Four parallel resolution streams exchange features at every stage; the full resolution stream is never discarded, so a two pixel bund survives to the head. ``` One caveat: HRNet's finest stream runs at a quarter of the input resolution, because the stem downsamples twice before the streams begin, so the logits are still upsampled four times at the head. ##### Preparing the imagery The client supplied high resolution optical imagery over the area of interest, plus parcels digitised by hand for part of it. Three preparation steps took longer than training did. Terrain first. Kishtwar is steep, and relief displacement in imagery not orthorectified against a decent elevation model shifts a hillside parcel by more than its own width relative to a polygon digitised from another acquisition. We orthorectified everything to one reference and checked alignment by overlaying the digitised parcels in QGIS before rasterising a label. Where references no longer sat on the terraces, we dropped those blocks rather than teach the network that boundaries live a few metres off the visible riser. Radiometry second. Scenes from different dates had different illumination, and terraces on the shaded side of the valley sat in a narrower dynamic range. We normalised each band per scene with robust percentiles rather than a global mean and standard deviation, so one bright rooftop could not compress the histogram. Splitting third. Random tile splits leak: a validation tile shares its edges, crops and lighting with the training tile next door. We split by contiguous geographic block, holding out whole blocks chosen to include both valley sides and both dense and sparse terracing. ##### Labels: a boundary channel, an interior mask and a distance weight The network predicts two channels. The first is the boundary: the digitised polygon edges rasterised into the image grid and dilated by a pixel or two, so the target is not a single pixel line the network can miss by one pixel and be scored as fully wrong. The second is the parcel interior, eroded by the same amount. We do not vectorise the interior channel, but it is an easier auxiliary task that shares features with the hard one, and it provides marker seeds for the watershed later. Boundary pixels are a tiny fraction of the image, so plain cross-entropy learns to predict background and stop. Dice loss on the boundary channel helps because it ignores the class ratio. What helped most was a per pixel weight map from the distance transform of the boundary labels: pixels near an edge carry more weight and pixels deep inside a large parcel very little, so the penalty concentrates at the transition. A boundary aware loss that penalises the distance between predicted and reference contours narrowed the edges further but was unstable early in training, when predictions were far from any contour; distance-weighted binary cross-entropy plus Dice was easier to tune and gave edges of the same quality once skeletonisation was in place. Augmentation was restricted on purpose. Flips, right-angle rotations and brightness jitter are useful; large scale augmentation is not, because the physical size of a terrace is diagnostic, and a network taught that parcels can be any size will split a large field along a plough line. ##### Tiling with overlap The rasters are far larger than any GPU input, so inference runs on tiles, and the problem with tiles is their edges. The network sees less context at a tile border, so predictions there are worse, and where two tiles meet the predictions disagree by a pixel or two. Vectorise that and every tile edge becomes a fake field boundary in a perfect grid across the map. We saw exactly that in the first run. The fix has two parts: read tiles with overlap so every pixel is predicted with full context on all sides, and blend the overlapping predictions so no pixel gets a hard hand-off from one tile to the next. The generator uses rasterio windows, so the whole raster is never loaded. ```python import numpy as np import rasterio import torch from rasterio.windows import Window def tile_windows(width, height, tile=512, overlap=128): stride = tile-overlap for r0 in range(0, max(height-overlap, 1), stride): for c0 in range(0, max(width-overlap, 1), stride): r1, c1 = min(r0+tile, height), min(c0+tile, width) yield Window(c0, r0, c1-c0, r1-r0) def ramp(n, overlap, eps=1e-3): w = np.ones(n, np.float32) k = min(overlap, n // 2) edge = 0.5*(1.0-np.cos(np.linspace(0.0, np.pi, k, endpoint=False))) w[:k] = edge w[n-k:] = edge[::-1] return np.maximum(w, eps) def tile_weight(h, w, overlap): return np.outer(ramp(h, overlap), ramp(w, overlap)) def pad32(x): h, w = x.shape[-2:] return np.pad(x, ((0, 0), (0, (-h) % 32), (0, (-w) % 32)), mode="edge"), (h, w) @torch.inference_mode() def predict_boundary(src, model, normalise, tile=512, overlap=128, device="cuda"): prob = np.zeros((src.height, src.width), np.float32) norm = np.zeros_like(prob) for win in tile_windows(src.width, src.height, tile, overlap): x, (h, w) = pad32(normalise(src.read(window=win))) logits = model(torch.from_numpy(x)[None].to(device)) p = torch.sigmoid(logits[0, 0]).cpu().numpy()[:h, :w] wgt = tile_weight(h, w, overlap) r, c = win.row_off, win.col_off prob[r:r+h, c:c+w] += p*wgt norm[r:r+h, c:c+w] += wgt return prob/norm ``` `ramp` is flat across the middle of the tile and falls along a raised cosine to a small floor across the overlap band. Where two tiles overlap, one weight rises as the other falls and they sum to one, so the blended probability moves smoothly between predictions. The floor keeps tiles on the raster's outer edge, which have no neighbour, from dividing by zero. `pad32` exists because HRNet's coarsest stream is at one thirty-second of the input; edge tiles not a multiple of 32 wide would fail inside the network. ```figure hrnet-kishtwar-boundaries/tiling A tile slides across the raster with an overlap band; cosine ramp weights hand off from one tile to the next so no seam survives into the vector layer. ``` We first tried cropping a halo off each tile and pasting only the core. That removes the worst context effects but leaves a faint step where the two predictions disagree across the crop line, and skeletonisation turned each step into a kink. Weighted blending removed them. The overlap must be at least as wide as the context the network uses at a boundary pixel, which for HRNet is large, so we erred on the generous side and paid in inference time. ##### From probability map to parcels The blended output is one region-sized raster of boundary probability. Turning it into instances is a morphology problem, and each step exists because a naive version failed. Thresholding first. A single threshold either breaks weak bunds into dashes or lets noise through. Hysteresis thresholding keeps a pixel if it exceeds a low threshold and connects to a pixel above a high one, which keeps long faint risers and drops speckle. Thinning next. The dilated labels produce predicted edges a few pixels wide; skeletonising them to one pixel gives clean lines. Skeletons come with spurs, short dangling branches where the edge was noisy, so we prune any short branch ending in a free endpoint. A small morphological closing before skeletonisation joins hairline gaps that would let neighbouring parcels leak into each other. Then instances. Our first attempt was connected components on the non-boundary pixels, and it fails in an ugly way: any single pixel gap in the skeleton, and there are always some under tree shadow, merges two parcels. Watershed behaves better. We take the distance transform of the interior mask, seed markers at its strong local maxima, and flood from them over the boundary probability as an elevation surface. A weak bund is still a ridge, so with a marker on each side the flood stops there. The remaining failure is the opposite one: two markers inside one large field split it along a plough line. We merge adjacent basins whose shared ridge is low in probability, a threshold on the saddle rather than the pixel. ##### Vectorisation and topology cleanup The label raster is polygonised with `rasterio.features.shapes`, which walks the pixel boundaries and produces one polygon per label. Those polygons are topologically perfect and visually awful: staircase edges at pixel resolution, thousands of vertices per parcel. Simplifying each polygon on its own destroys the topology, because the simplifier moves a shared edge differently for the two parcels that share it, leaving slivers and overlaps along it. The fix is to simplify the shared edges once, as a network of arcs, and rebuild the faces afterwards. ```python import numpy as np import geopandas as gpd import shapely from rasterio import features from shapely.geometry import shape from shapely.ops import polygonize, unary_union def label_raster_to_parcels(labels, transform, crs, grid=0.05, tol=0.3, min_area=25.0): faces = [shape(g) for g, _ in features.shapes(labels.astype(np.int32), mask=labels > 0, transform=transform, connectivity=4)] rings = [f.exterior for f in faces] + [r for f in faces for r in f.interiors] arcs = unary_union(rings) arcs = shapely.set_precision(arcs, grid) arcs = shapely.simplify(arcs, tol, preserve_topology=True) parcels = gpd.GeoDataFrame(geometry=list(polygonize(arcs)), crs=crs) return merge_slivers(parcels, min_area) def merge_slivers(gdf, min_area): geoms = gdf.geometry.tolist() sindex = gdf.sindex for i, g in enumerate(geoms): if g is None or g.area >= min_area: continue near = [j for j in sindex.query(g, predicate="touches") if j != i and geoms[j] is not None] if not near: continue j = max(near, key=lambda k: g.intersection(geoms[k]).length) geoms[j] = unary_union([geoms[j], g]) geoms[i] = None return gpd.GeoDataFrame(geometry=[g for g in geoms if g is not None], crs=gdf.crs) ``` `unary_union` over all the rings nodes the linework: every junction of three parcels becomes a vertex, and every stretch between junctions becomes a single arc that both neighbours reference. `simplify` never moves a line's endpoints, so junctions stay fixed and each shared arc is simplified exactly once. `set_precision` snaps coordinates to a grid first, so vertices meant to coincide become identical rather than a near miss. `polygonize` rebuilds a face for every closed ring; dangling arcs close no ring and produce no face, which disposes of them for free. Faces below a minimum area are slivers and are merged into the neighbour sharing the longest edge. All of this runs in the raster's projected CRS so tolerances are in metres; the result is reprojected to WGS84 at the end, because GeoJSON requires it. ```figure hrnet-kishtwar-boundaries/vectorise From boundary probability to thinned skeleton, to watershed instances, to simplified polygons that still share every edge exactly. ``` ```mermaid flowchart LR IMG[Orthorectified imagery] --> TILE[Overlap tiles] TILE --> NET[HRNet-W48] NET --> BLEND[Weighted stitch] BLEND --> MORPH[Hysteresis, thin, prune] MORPH --> WS[Watershed] WS --> VEC[Arcs, simplify, polygonize] VEC --> QA[Overlay review in QGIS] QA --> GIS[GeoJSON and GeoPackage] QA -- corrections --> NET ``` ##### QA against the digitised parcels We assessed the output qualitatively by overlaying it on the manually digitised parcels in QGIS with the imagery underneath. The aim was not a score but to name the kinds of disagreement, so we knew which to fix in the model and which to leave to reviewers. Four recurred: missed risers under tree shadow, where the image contains no edge; merged strips where a grass covered bund was indistinguishable from the fields either side; invented edges along footpaths and drainage lines, which look exactly like bunds and sometimes are; and, in the earliest runs, the tile grid, which is how we found the seam problem. A share of the disagreements were errors in the reference, where the digitiser had drawn through a terrace since split or missed a narrow strip; those were corrected and fed back as training labels. Each reviewed polygon carries a review flag, and a confidence attribute, the mean boundary probability along its edge, lets the client's team sort to the doubtful ones. ##### Delivery to the client's GIS The deliverable is a GeoJSON parcel layer in EPSG:4326 with a stable parcel identifier, area in square metres computed in the projected CRS, the block identifier, the confidence and review flags, and the model version. Because GeoJSON has no spatial index and a region's parcels make a large file, the same layer also ships as a GeoPackage, plus the stitched probability raster as a Cloud Optimised GeoTIFF for auditing individual edges. A QGIS style file colours parcels by confidence, so a reviewer sees first where to look. ##### What we would change Next time the head would predict at full resolution rather than one quarter, at the cost of memory. We would also label terrace risers as a class separate from bunds on flat land, because the network learned them as two visually different things and the single boundary class made it hesitate at the transition between them. ### Crop Type Classification from Sentinel-2 Time Series - Stack: Sentinel-2 L2A, xarray, PyTorch, XGBoost, PostGIS - Summary: Per parcel crop mapping for Zaraat Dost Private Limited: cloud-masked Sentinel-2 index series resampled onto a regular grid, a bidirectional GRU with attention pooling, and scheduled district-wide inference delivered as a PostGIS layer. - Challenge: Label every parcel in a district with its crop from Sentinel-2 alone, in a region where the monsoon removes weeks of usable imagery, the scene classification layer misses haze and small shadows, survey labels carry spatial and temporal errors, and two major crops dwarf the minor crops the client most wanted mapped. - Solution: We built per-parcel index series from L2A with a combined SCL and U-Net mask, resampled them onto a ten-day grid with explicit observed and age channels instead of smoothing over gaps, moved from an XGBoost baseline through a temporal CNN to a bidirectional GRU with attention pooling, validated with spatial block folds and a held-out season, and ran inference as idempotent scheduled stages that write a per-parcel prediction layer to PostGIS. #### Crop type classification from Sentinel-2 time series Zaraat Dost Private Limited needed to know what was growing in every parcel of a district without sending a surveyor to every field. Sentinel-2 passes over the same ground every five days, and a crop reveals itself less by how it looks on one date than by how its greenness rises and falls across the season. So this was a time series problem wearing a satellite costume: turn a stack of cloud-riddled scenes into one clean sequence per parcel, learn the shape of each crop's season, and put a label with a confidence on every polygon in the district. The parcel polygons already existed from the farm digitisation work, so the unit of prediction was one label per parcel, not per pixel. The crop calendar has two seasons, rabi in winter and kharif in summer, with the monsoon in the middle of kharif, and the client wanted the map refreshed as new imagery arrived and delivered as a layer their agronomists could open in the tools they already used. ##### Building per parcel time series from L2A We worked from Level-2A products, which are atmospherically corrected and ship with a scene classification layer (SCL). From the thirteen bands we kept the four 10 m bands (blue, green, red, near infrared), the red-edge and narrow near infrared bands at 20 m, and the two shortwave infrared bands, with the 20 m bands resampled onto the 10 m grid by nearest-neighbour so no sub-pixel structure was invented. From those bands we computed a small set of indices per pixel per date: NDVI for overall greenness, EVI because it saturates later over dense canopy, NDRE from the red edge because it separates crops at a stage where NDVI has already plateaued, and NDMI from near infrared and shortwave infrared because it tracks canopy water and picks out flooded rice fields early in kharif. Raw bands went into early experiments and came out again: they generalised worse across seasons, because reflectance shifts with sun angle and atmosphere in ways the ratios largely cancel. Each index was reduced to one number per parcel per date by taking the median of the pixels inside the polygon, with a one-pixel inward buffer so mixed edge pixels did not drag the value toward the neighbouring field. The median mattered: a small parcel with one shadowed pixel or a track running through it would have had its mean pulled down badly. ##### Cloud and shadow masking The SCL flags clouds, cloud shadows, cirrus, snow and saturated pixels, and it is a reasonable first filter. Used alone, it let through two kinds of contamination that wrecked our early sequences. Thin haze at cloud edges passed as clear and produced NDVI values slightly below the truth, which looks exactly like a crop under stress. Small cumulus shadows over a single field were missed entirely, because the SCL runs at 20 m and a small shadow hides inside one "clear" pixel. So we combined the SCL with our own cloud and shadow segmentation model, a U-Net on multispectral input that we had built for exactly this purpose, and treated a pixel as bad if either source said so. We then dilated the bad region by a couple of pixels so cloud edges and their haze were dropped too. The last check happens at the parcel level: if too few of a parcel's pixels were clear on a date, that observation was dropped rather than trusted. ```python import numpy as np import pandas as pd import xarray as xr CLOUDY_SCL = np.array([0, 1, 3, 8, 9, 10, 11]) # nodata, saturated, shadow, cloud med/high, cirrus, snow def valid_mask(scl: xr.DataArray, custom: xr.DataArray, dilate: int = 2) -> xr.DataArray: """True where a pixel is usable. scl is the L2A scene classification, custom is the U-Net cloud and shadow probability on the same grid.""" bad = scl.isin(CLOUDY_SCL) | (custom > 0.4) if dilate: k = 2 * dilate + 1 bad = bad.astype("uint8").rolling(y=k, x=k, center=True, min_periods=1).max() > 0 return ~bad def parcel_series(ndvi: xr.DataArray, valid: xr.DataArray, parcel_id: xr.DataArray, min_valid_frac: float = 0.7) -> pd.DataFrame: """Median NDVI per parcel per acquisition, dropped when too few pixels are clear.""" masked = ndvi.where(valid) frames = [] for t in ndvi.time.values: df = pd.DataFrame({ "parcel": parcel_id.values.ravel(), "ndvi": masked.sel(time=t).values.ravel(), "valid": valid.sel(time=t).values.ravel(), }) df = df[df.parcel > 0] g = df.groupby("parcel") out = g.ndvi.median().to_frame("ndvi") out["valid_frac"] = g.valid.mean() out["date"] = pd.Timestamp(t) frames.append(out.reset_index()) obs = pd.concat(frames, ignore_index=True) return obs[obs.valid_frac >= min_valid_frac].drop(columns="valid_frac") ``` ```figure sentinel2-crop-timeseries/masking-resampling Irregular L2A observations with cloud and shadow dates masked out, then resampled onto a regular ten-day grid with fills flagged explicitly. ``` ##### Gap filling and resampling to a regular grid After masking, each parcel had a ragged list of clear observations at irregular dates: five-day spacing where orbit tracks do not overlap, closer where they do, and long holes wherever weather had intervened. Sequence models want a fixed-length tensor, so we resampled every parcel onto the same grid of ten day steps across the season. Linear interpolation in time carried values across short gaps. The monsoon produces gaps of several weeks over the whole district, and linear interpolation across a six-week hole draws a confident straight line through exactly the period where rice is transplanted and greens up. We tried a Whittaker smoother and a Savitzky-Golay filter. They produced beautiful curves and hid the problem: the model could no longer tell an observed value from an invented one, and it learned to trust the invented ones. What worked was to keep the interpolation simple and make the uncertainty explicit. Each step carries three channels per index: the value, a flag saying whether a real clear observation fell within half a step of the grid date, and the age of the most recent clear observation, capped and scaled. Beyond a maximum gap we stop interpolating, set the value to zero, and let the flags say so. The network learns how much to discount a filled step, and the monsoon gap becomes a pattern it recognises rather than a lie it believes. ```python def resample_parcel(obs: pd.DataFrame, grid: pd.DatetimeIndex, max_gap_days: int = 40) -> np.ndarray: """obs holds one parcel's clear observations (columns date, ndvi). Returns an array of shape (T, 3): value, observed flag, scaled age of last observation.""" s = obs.set_index("date")["ndvi"].sort_index() s = s[~s.index.duplicated(keep="last")] # interpolate on the union of real dates and grid dates, then keep only grid rows filled = s.reindex(s.index.union(grid)).interpolate(method="time", limit_area="inside") value = filled.reindex(grid).to_numpy(dtype=float) stamps = pd.Series(s.index, index=s.index) nearest = stamps.reindex(grid, method="nearest", tolerance=pd.Timedelta(days=5)) observed = nearest.notna().to_numpy() last = stamps.reindex(grid, method="ffill").reset_index(drop=True) age = pd.Series(grid).sub(last).dt.days.to_numpy(dtype=float) # NaN before the first observation age = np.nan_to_num(age, nan=max_gap_days) stale = (age > max_gap_days) | np.isnan(value) value[stale] = 0.0 # zeroed but flagged, so the network learns to ignore rather than trust it return np.stack([value, observed.astype(float), np.minimum(age, max_gap_days) / max_gap_days], axis=1) ``` The same function runs for every index and the arrays are concatenated along the channel axis. ##### The label problem Ground truth came from field surveys: teams walking parcels with a GPS app and recording the crop. The most common error was spatial: a surveyor standing on a bund between two fields would tap the wrong polygon, so the label described the neighbour. The second was temporal: surveys spread across weeks, and a field recorded as fallow one week might be sown the next, so label and imagery described different states. Intercropped parcels were eventually excluded because no single label was true. Class imbalance was severe. Two major crops covered the large majority of surveyed parcels, and some minor crops had a few dozen examples in total. Plain cross-entropy produced a model that predicted the majority class, looked fine on overall accuracy, and was useless for the minor crops the client cared most about. We merged the rarest classes into an "other" bucket for the first release, weighted the loss by inverse class frequency with a square-root taper so the weights did not explode, and oversampled minor-crop parcels within each epoch. We also turned the model on its own labels: parcels where it disagreed with the survey at high confidence went back to the survey team, and a meaningful share turned out to be survey errors. ##### Models, in the order we tried them The baseline was XGBoost on hand-crafted temporal features: for each index the maximum, minimum, mean, amplitude, the step at which the maximum occurred, the slopes of green up and senescence, the integral over the season, plus the raw values at each grid step. It trained in minutes and was strong on the major crops. It struggled on the minor crops and, more seriously, it degraded on a season where sowing had been late, because features like "step of maximum" move with the calendar and trees split on exact thresholds. The second model was a one-dimensional temporal CNN: a few convolutional layers over the time axis with small kernels, batch normalisation, and global average pooling into a linear head. Convolution is shift-equivariant, so a late season looks like the same pattern moved along the axis, and the CNN handled calendar shift far better than the trees. The pooling threw away where in the season things happened, though, and two crops with similar green up shapes at different dates became confusable. The third was a bidirectional GRU over the resampled sequence, with the observed and age channels as input. We tried an LSTM first and switched to the GRU because it trained faster with no loss we could detect at this sequence length. Reading in both directions means the state at each step knows what came before and after, which is what you need to say "this dip in the middle is a gap, not a harvest." On top of the recurrent states we placed an attention pooling layer that learns a weight per step and takes a weighted sum, instead of using only the final hidden state. That single change made the largest difference on minor crops, and the attention weights became a diagnostic: for rice they concentrated on the transplanting flood and the late peak, for wheat on the winter green up. ```figure sentinel2-crop-timeseries/temporal-model The bidirectional GRU reads the ten-day sequence in both directions, attention pooling weights the informative steps, and the head emits per-class confidence. ``` The GRU with attention became the final model. It won on the metrics we trusted (spatially blocked validation and a held out season), it gave calibrated confidences after temperature scaling, and it needed none of the feature engineering the baseline did. A transformer encoder was tried briefly and dropped, because with our label count it overfitted quickly and gave nothing the GRU did not. ```python import torch from torch import nn class TemporalGRUClassifier(nn.Module): """Input (B, T, C): T grid steps, C = indices x (value, observed, age). Output (B, n_classes) logits.""" def __init__(self, in_ch: int, n_classes: int, hidden: int = 64, layers: int = 2, dropout: float = 0.2): super().__init__() self.inp = nn.Sequential(nn.Linear(in_ch, hidden), nn.GELU()) self.gru = nn.GRU(hidden, hidden, num_layers=layers, batch_first=True, bidirectional=True, dropout=dropout if layers > 1 else 0.0) self.attn = nn.Sequential(nn.Linear(2 * hidden, hidden), nn.Tanh(), nn.Linear(hidden, 1)) self.head = nn.Sequential(nn.Dropout(dropout), nn.Linear(2 * hidden, n_classes)) def forward(self, x: torch.Tensor, return_attention: bool = False): h, _ = self.gru(self.inp(x)) # (B, T, 2H), each step sees past and future w = torch.softmax(self.attn(h).squeeze(-1), dim=1) # (B, T), one weight per step ctx = torch.einsum("bt,bth->bh", w, h) # weighted sum over time logits = self.head(ctx) return (logits, w) if return_attention else logits ``` ##### Training without fooling ourselves Two kinds of leakage make a crop model look better than it is. The first is spatial. Neighbouring parcels share soil, water and sowing dates, so a random split puts near-duplicates of the test parcels into training and inflates the score. We split by spatial block instead: the district was cut into a grid of square blocks, every parcel was assigned to the block containing its centroid, and blocks were dealt into folds. A test parcel then never had a training neighbour in its own block, and parcels within a buffer of a fold boundary were dropped from training to limit leakage across the edge. ```figure sentinel2-crop-timeseries/spatial-blocks Validation folds are dealt by spatial block so a test parcel never has a training neighbour, and one whole season is held out to measure transfer to a different calendar. ``` The second leakage is temporal. If a parcel appears in training with one season and in test with the next, the model can learn that this field is always rice. We kept whole seasons separate: train and validate within the seasons we had, then measure on a season that was never seen. The held out season scores were consistently lower than the blocked cross-validation scores, and closing that gap became the actual job. Augmentation was done at the observation level rather than on the tensors. For each training parcel we randomly dropped a fraction of the clear observations, shifted all dates by a few days to mimic a late or early season, scaled the amplitude slightly, and then ran the same resampler as production. Jittering the resampled tensor directly had produced fills that never occur in real data. ```python def augment_observations(obs: pd.DataFrame, rng: np.random.Generator, max_shift_days: int = 6, drop_p: float = 0.15) -> pd.DataFrame: """Drop random clear observations and shift the season, then resample as usual.""" out = obs[rng.random(len(obs)) >= drop_p].copy() shift = int(rng.integers(-max_shift_days, max_shift_days + 1)) out["date"] = out["date"] + pd.to_timedelta(shift, unit="D") out["ndvi"] = out["ndvi"] * rng.uniform(0.95, 1.05) return out ``` ##### Inference at district scale The client wanted the map refreshed as the season progressed, over a district's worth of parcels, so the pipeline was written as scheduled, idempotent stages. A scheduler wakes on a fixed cadence, lists new L2A tiles intersecting the district, and records them in a manifest. A masking stage runs the SCL and U-Net masks per tile and writes valid-pixel rasters. An aggregation stage appends new dates to a per parcel observation table. A resampling stage rebuilds the grid tensors. The model stage batches parcels through the GRU on one GPU and writes a prediction row per parcel per run: class, confidence, number of clear observations so far, and run date. ```mermaid flowchart LR S[Scheduler] --> M[New L2A tiles into manifest] M --> K[Mask: SCL + U-Net] K --> A[Per-parcel medians appended] A --> R[Resample to 10-day grid] R --> G[GRU inference in batches] G --> P[(PostGIS predictions)] P --> L[Vector tiles and GeoJSON] ``` Every stage checks the manifest before doing work, so a rerun after a failure skips finished tiles and never leaves the table in a mixed state. Early in the season the confidences are low because the sequence is mostly zeros and flags, and they rise as each crop's distinguishing phase arrives; the per-run history shows when the model made up its mind, and a parcel whose label flips late usually has a problem. ##### Delivering the layer Predictions landed in a PostGIS table keyed by parcel id, joined to the parcel polygons, and were served as vector tiles and as GeoJSON exports for QGIS. Each feature carries the class, the confidence, the count of clear observations and the date of the last run. Parcels below a confidence threshold are styled as unsure rather than forced into a class, and a filter for them lets survey teams go where the model needs help. Confirmed survey results flow back into the label set for the next training round. ##### What we would do first next time The mask and the resampler, not the model. Every model trained on the earlier, leakier sequences looked fine in a random split and fell apart on a held out season, and the fix was upstream every time. ### Serving Remote Sensing Models: On Demand Inference and Scheduled Jobs - Stack: Docker, FastAPI, ONNX Runtime, PostgreSQL, Cron / Airflow, MLflow - Summary: The deployment layer behind the farmland models: ONNX export with a parity gate, a FastAPI service with a warm session and bounded concurrency, scheduled jobs that claim work by idempotency key and write versioned rows to PostGIS, and an MLflow promotion flow that rolls back by moving an alias. - Challenge: The segmentation and classification models lived as PyTorch checkpoints run by hand, so every new area of interest meant a person, a machine and a script. Serving them needed a runtime that does not silently change the model, a service that stays responsive under concurrent requests, scheduled runs that survive duplicate ticks and evicted containers, and a way to know which model version produced any row in the database. - Solution: Models are exported to ONNX with dynamic axes and only registered after a parity check on held-out tiles; a FastAPI service holds one warm ONNX Runtime session per process behind a semaphore that rejects rather than queues. Scheduled jobs discover new scenes with a spatial query, claim each unit through a unique idempotency key, retry with backoff, and write outputs and run status in one transaction. MLflow aliases decide which version runs, every output carries that version, drift on input statistics is flagged on the run row, and rollback is the promotion procedure in reverse. #### Serving remote sensing models: on demand inference and scheduled jobs The segmentation and classification models we built for farmland mapping spent their first months as PyTorch checkpoints and notebooks. Every request to run them over a new area meant someone opening a machine, loading weights and running a script by hand. This case study covers the deployment layer that replaced that: containerised FastAPI services for on demand inference, scheduled jobs that pull new imagery and write results to PostGIS, and the versioning, logging and drift checks around them. The stack is Docker, FastAPI, ONNX Runtime, PostgreSQL with PostGIS, cron and Airflow for scheduling, and MLflow for the model registry. ##### Exporting to ONNX and proving parity We serve ONNX rather than PyTorch for two reasons. The inference image does not need the training framework, so it is smaller and starts faster, and ONNX Runtime gives us a stable, well-tested CPU path with explicit threading control. The cost is an export step that can silently change what the model computes, so the export is never trusted without a parity check. Two classes of model go through it. The segmentation models take a multi-band tile and return a per pixel map; the classification models take a stack of Sentinel-2 observations for a parcel and return class logits. Both are exported with a dynamic batch axis, and the segmentation models also with dynamic height and width so tile size is a deployment choice, not a training-time constant. The parity check runs the original PyTorch model and the exported graph on the same held out inputs and compares outputs under an absolute and relative tolerance. It also compares the argmax maps, because a small numerical difference in logits is acceptable while a flipped class at a boundary pixel is exactly the kind of regression we need to see. ```python import numpy as np import onnxruntime as ort import torch def export_and_check(model, sample: torch.Tensor, path: str, opset: int = 17, atol: float = 1e-4, rtol: float = 1e-3) -> dict: model.eval() with torch.no_grad(): reference = model(sample).cpu().numpy() torch.onnx.export( model, sample, path, opset_version=opset, input_names=["x"], output_names=["logits"], dynamic_axes={"x": {0: "batch", 2: "height", 3: "width"}, "logits": {0: "batch", 2: "height", 3: "width"}}, do_constant_folding=True, ) sess = ort.InferenceSession(path, providers=["CPUExecutionProvider"]) got = sess.run(["logits"], {"x": sample.cpu().numpy()})[0] close = np.allclose(got, reference, atol=atol, rtol=rtol) argmax_agree = float((got.argmax(1) == reference.argmax(1)).mean()) max_abs = float(np.abs(np.subtract(got, reference)).max()) ok = close and argmax_agree >= 0.999 if not ok: raise RuntimeError(f"parity failed: max_abs={max_abs:.2e} argmax_agree={argmax_agree:.4f}") return {"max_abs_diff": max_abs, "argmax_agreement": argmax_agree, "opset": opset} ``` The first exports failed this gate and the reasons were instructive. A bilinear upsampling layer with `align_corners` set the way the training code had it exported to a resize whose corner handling differed by a pixel, which shifted the whole boundary map by half a pixel and showed up as an argmax agreement in the high nineties rather than near one. A model that had been left in training mode exported its batch normalisation with batch statistics. A custom padding step used a Python integer computed from the input shape, which the tracer baked in as a constant, so the model worked at the export tile size and produced garbage at any other. Each of these would have shipped without the check, and each is now a line in the export notes. ```figure model-serving-scheduled-inference/export-parity The checkpoint is exported with dynamic axes and run side by side with the original on held-out tiles; the gate compares logits under tolerance and argmax maps for agreement before the artifact is registered. ``` The parity result is logged as metrics on the MLflow run that registers the artifact, so the registry entry carries its own evidence. ##### ONNX Runtime session options A serving process holds exactly one session per model, created at startup, and never creates sessions per request. Session creation is the expensive part: the graph is optimised, weights are laid out, and thread pools are allocated. We set the graph optimisation level to the full extended set, and we serialise the optimised graph to disk once so container restarts skip that work. Threading is where the defaults are wrong for a service. ONNX Runtime will happily use every core for a single request through its intra-op pool, which is ideal for a batch script and poor for a service handling several requests at once, where the requests then fight over the same cores. We pin the intra-op thread count to a small number, disable inter-op parallelism because our graphs are sequential, and let the concurrency limiter in the service decide how many requests run at once. The product of the two is kept at or below the core count of the container. Tile inputs are pre-allocated in the right layout and passed with input and output binding, so the runtime does not copy arrays on the way in or out. For the segmentation models this was the single largest latency improvement after threading. ##### The FastAPI service The on demand service answers a small set of questions: run this model over this area, or over this uploaded tile, and return the result. Requests are validated with Pydantic before any work happens. A bounding box has to be in a known CRS, its area is capped, tile size is restricted to the values the model was validated on, and the model is named by registry alias, never by file path. Bad requests fail fast with a message that says which field was wrong. ```python import asyncio from contextlib import asynccontextmanager import numpy as np import onnxruntime as ort from fastapi import FastAPI, HTTPException from fastapi.concurrency import run_in_threadpool MAX_INFLIGHT = 2 @asynccontextmanager async def lifespan(app: FastAPI): opts = ort.SessionOptions() opts.intra_op_num_threads = 4 opts.execution_mode = ort.ExecutionMode.ORT_SEQUENTIAL opts.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_EXTENDED app.state.session = ort.InferenceSession(resolve_model_path("segmentation@production"), opts) app.state.slots = asyncio.Semaphore(MAX_INFLIGHT) warm = np.zeros((1, 4, 512, 512), dtype=np.float32) app.state.session.run(None, {"x": warm}) yield app = FastAPI(lifespan=lifespan) @app.post("/infer/tile") async def infer_tile(req: TileRequest): slots: asyncio.Semaphore = app.state.slots if slots.locked(): raise HTTPException(status_code=503, detail="inference capacity exhausted, retry later") async with slots: x = decode_tile(req) logits = await run_in_threadpool(app.state.session.run, ["logits"], {"x": x}) return encode_result(logits[0], req) ``` Two decisions in that handler are worth defending. The warm-up call in `lifespan` runs one full-size input through the session before the service reports healthy, because the first run of a session pays for lazy allocations and is several times slower than the steady state; without it the first real request after every deploy was the slowest of the day. And the semaphore rejects when full rather than queueing. A queue hides saturation: latency climbs quietly until a client times out anyway. A 503 with a retry hint lets the caller back off and lets the orchestrator scale, and it keeps the memory of the container bounded, since each in-flight segmentation request holds tiles and logits that are not small. ```figure model-serving-scheduled-inference/request-path A request is validated, takes one of a fixed number of slots, runs through the warm session in a worker thread and returns; when every slot is taken the service answers 503 instead of queueing. ``` The health endpoint reports the loaded model version and the parity metrics it was registered with, which turned out to be the fastest way to answer the question "what is actually running" during an incident. ##### Scheduled jobs Most inference is not on demand. It is the routine of new imagery arriving and results needing to exist before someone asks. A scheduled job wakes on a cron expression, or as an Airflow DAG where a workflow has several dependent steps, and does the following: discover scenes that have appeared since the last run and intersect the areas of interest, decide which scene and area pairs still need processing, tile and run the model, and write the outputs to PostGIS with the model version attached to every row. Discovery is a query, not a crawl. The imagery index is a table of scenes with a footprint geometry, acquisition time and cloud fraction, kept up to date by its own ingestion job. The inference job intersects it with the areas of interest table in one spatial query and takes what is new. Airflow was worth its weight for the workflows with multiple steps, for example masking clouds and shadows before classification, because it gives per-task retries and a visible history. Plain cron remained for the single-step jobs, where a scheduler with a web interface and a metadata database would have been more infrastructure than the job. ##### Idempotent runs and retries Scheduled work fails in dull ways: a network hiccup halfway through a scene, a container evicted, the same tick firing twice after a scheduler restart. The rule that made this survivable is that every unit of work has an idempotency key, and doing the same unit twice is harmless. The key is a hash of the things that define an output: the model name and registered version, the scene identifier, the area of interest identifier, and the parameters that affect the result, such as tile size and overlap. A `runs` table with a unique constraint on that key is the source of truth. A run is inserted as `pending` before any work happens; if the insert conflicts, the work is already done or in progress and the job skips it. Outputs are written in a transaction together with the update that marks the run `done`, so a crash between writing polygons and marking the run leaves nothing half-committed. ```python import hashlib import json import time import psycopg def idempotency_key(model: str, version: int, scene_id: str, aoi_id: int, params: dict) -> str: payload = json.dumps( {"model": model, "version": version, "scene": scene_id, "aoi": aoi_id, "params": params}, sort_keys=True, ) return hashlib.sha256(payload.encode()).hexdigest() def claim_run(conn: psycopg.Connection, key: str, meta: dict) -> bool: with conn.cursor() as cur: cur.execute( """ insert into runs (idem_key, model, model_version, scene_id, aoi_id, params, status) values (%(key)s, %(model)s, %(version)s, %(scene)s, %(aoi)s, %(params)s, 'pending') on conflict (idem_key) do nothing """, {"key": key, **meta}, ) return cur.rowcount == 1 def run_unit(conn: psycopg.Connection, key: str, meta: dict, work, attempts: int = 4) -> None: if not claim_run(conn, key, meta): return delay = 30.0 for attempt in range(1, attempts + 1): try: with conn.transaction(): rows = work() write_outputs(conn, rows, key) conn.execute("update runs set status = 'done', finished_at = now() where idem_key = %s", (key,)) return except Exception as exc: conn.execute( "update runs set status = 'retrying', attempts = %s, last_error = %s where idem_key = %s", (attempt, str(exc)[:500], key), ) if attempt == attempts: conn.execute("update runs set status = 'failed' where idem_key = %s", (key,)) raise time.sleep(delay) delay *= 2 ``` Retries back off exponentially and are capped, and the error text is kept on the run row. The `runs` table is also the run log the project entry mentions: which model version processed which scene, when, how many attempts it took, and what it said when it failed. A `pending` row older than the longest plausible run is treated as abandoned and reclaimed by the next tick, which handles the evicted container case without a separate reaper. ##### Model versioning and promotion with MLflow Every trained model is logged to MLflow with its parameters, evaluation metrics, the parity results from export, and the ONNX artifact, and registered under a model name. The services and the jobs never refer to a file. They refer to a registry alias, `production` or `candidate`, and resolve it to a specific version at startup. That version number is written into every output row and into every run. Promotion is a checklist, not a button. A candidate must pass the parity gate, must meet or beat the current production version on the held out evaluation set, and must run in shadow on the last few scheduled ticks, producing outputs into a separate schema so the two versions can be compared over real recent imagery rather than the training distribution. Only then is the alias moved, and moving the alias is a one-line change that is itself logged. The services pick up the new version on their next restart, which is a deliberate choice: hot-swapping a session inside a live process is possible, but a restart is a clean boundary and the containers start in seconds. ```figure model-serving-scheduled-inference/job-loop Each tick discovers new scenes over the areas of interest, claims each unit by idempotency key, runs the model resolved from the registry alias and writes versioned rows to PostGIS; promotion moves the alias, rollback moves it back. ``` ##### Monitoring drift on input statistics We do not have labels for new imagery as it arrives, so output quality cannot be measured directly. What can be measured is the input. For every processed scene the job records per-band mean and standard deviation, the fraction of pixels masked as cloud or shadow, and a coarse histogram, and compares them against the same statistics on the training data. A large distance on a band, or a masked fraction far above the seasonal norm, is flagged on the run row and raised as an alert. The check flags; it does not block. A monsoon scene is far from the training distribution and still needs processing, and a downstream analyst can filter by the flag. The value of the check has been in catching the boring failures early: a change in the imagery provider's scaling that shifted every band by a constant, and a masking job that silently stopped and left clouds in the inputs, both visible in the statistics before anyone looked at the polygons. ##### Rollback Because every output carries its model version and every process resolves an alias at startup, rollback is the promotion procedure run backwards. The alias is pointed at the previous version, the services restart, and the next scheduled tick uses the old model. Outputs produced by the bad version are not deleted; they are identifiable by version and can be reprocessed, because their idempotency keys include the version and a rerun with the restored version is a new unit of work. The shadow schema from the promotion step also makes the comparison available after the fact, so the decision to roll back is made on the same evidence that supported the promotion. ### An Enterprise RAG Pipeline That Can Say It Does Not Know - Stack: pgvector, PostgreSQL, FastAPI, Redis, Claude API - Summary: Grounded question answering over a private knowledge base, built so that refusal is a designed outcome: structure aware chunking, hybrid retrieval fused by reciprocal rank, a cross-encoder reranker, verified citations, and an evaluation harness that existed before any prompt was tuned. - Challenge: The client wanted question answering over policies, manuals and support tickets, with one non-negotiable: a confident wrong answer about a policy or a safety procedure was worse than no answer. Fixed-window chunking split tables and lost section context, dense retrieval missed exact product codes, and without a measurement of retrieval and faithfulness there was no honest way to tune anything. - Solution: Documents are parsed into trees and chunked along their own structure with heading paths attached. Dense pgvector search and BM25 run side by side and are merged with reciprocal rank fusion, a cross-encoder reranks the candidates, and a gate compares the best score with a calibrated threshold before generation. Every cited sentence is verified against its source span, unsupported answers are rejected, and a CI harness reports retrieval recall at k, faithfulness and refusal accuracy for every release, with content-hash, query and reranker caches keyed by index version. #### An enterprise RAG pipeline that can say it does not know The client had a private knowledge base of policies, procedures, product manuals and resolved support tickets, and wanted a question answering system over it. The requirement that shaped the whole design arrived in the first meeting: a confident wrong answer about a leave policy or a safety procedure was worse than no answer at all. So the system had to cite what it used, and it had to refuse when it had nothing worth citing. Most of the engineering below exists to make refusal a first-class outcome rather than an accident. ##### Ingestion and structure aware chunking The first version chunked documents into fixed windows of a few hundred tokens with some overlap, because that is what every tutorial does. It failed in ways that were obvious once we looked at retrieved passages. Tables were cut across rows, so a chunk might contain the header of a rate table and none of the rates. Numbered procedure steps were split so that step four appeared without the warning in step three. Worst of all, a chunk from deep inside a document had no idea which document or section it came from, so a passage that stated a day limit retrieved well for the wrong policy. The replacement parses every document into a tree first. PDFs go through a layout-aware parser that recovers headings, paragraphs, lists and tables; Word documents and Markdown already carry that structure; HTML is cleaned to the same shape. Chunking then walks the tree and packs sibling paragraphs under one heading into a chunk until a token budget is reached. Tables are never split: a table that exceeds the budget becomes its own chunk with its caption, and if it is still too large it is split by row groups with the header row repeated. Lists are kept with their introductory sentence. Every chunk is prefixed with its heading path, for example "Leave policy › Annual leave › Carry-over", so that both the embedding and the keyword index see the context the passage lives in. We also record for every chunk the document identifier, the heading path, character offsets into the source and a content hash. The offsets are what make citations point at a highlightable span later. The hash is what makes re-ingestion cheap, which mattered more than we expected. ```figure enterprise-rag-pipeline/chunking A parsed document tree is packed into chunks along its own structure, each prefixed with its heading path and written to a dense index and a keyword index. ``` ##### Two indexes, because neither is enough Dense retrieval is good at meaning and bad at exact strings. Keyword retrieval is the reverse. An enterprise corpus is full of exact strings that matter: product codes, form numbers, error identifiers, the names of internal systems. A user asking about "form HR-27" should get the chunk that contains "HR-27", and an embedding model has no reliable opinion about that token. Chunks are embedded and stored in PostgreSQL with pgvector, using an HNSW index with cosine distance. Alongside, the same chunk text is tokenised and stored for BM25 scoring. We first tried PostgreSQL's built-in full-text ranking and found it under-weighted rare terms relative to BM25, which is exactly the property we needed for codes and identifiers, so we moved to a proper BM25 scorer over the same tokens. Both indexes are versioned together: an index version is stamped on every chunk, and a query only ever runs against one version. ##### Hybrid retrieval with reciprocal rank fusion Each query is run against both indexes and each returns a ranked list. Merging them by score is a mistake we made once: cosine similarities and BM25 scores live on different scales, and any linear combination is tuned to one corpus and breaks on the next. Reciprocal rank fusion ignores scores entirely and rewards documents that appear high in either list, with a bonus for appearing in both. It has one constant, and it is robust to that constant. ```ts export interface Ranked { chunkId: string; rank: number } export function reciprocalRankFusion( lists: Ranked[][], k = 60, limit = 40, ): { chunkId: string; score: number }[] { const scores = new Map(); for (const list of lists) { for (const { chunkId, rank } of list) { const prev = scores.get(chunkId) ?? 0; scores.set(chunkId, prev + 1 / (k + rank)); } } return [...scores.entries()] .map(([chunkId, score]) => ({ chunkId, score })) .sort((a, b) => (b.score > a.score ? 1 : b.score < a.score ? -1 : a.chunkId.localeCompare(b.chunkId))) .slice(0, limit); } ``` Ranks here are one-based. The tie-break on chunk identifier looks pedantic, but it makes the fused list deterministic, and deterministic retrieval is what allows the evaluation harness to compare two releases without noise. The behaviour we care about most is that a chunk sitting in the middle of both lists can rise above a chunk that tops one list and is absent from the other, which is what consensus between two different notions of relevance should mean. ```figure enterprise-rag-pipeline/hybrid-fusion Dense and keyword results are fused by reciprocal rank: a chunk ranked in the middle of both lists rises above one that only one retriever liked. ``` ##### Reranking The fused list is a candidate set, not an answer set. The top forty candidates go to a cross-encoder reranker that reads the query and each chunk together and produces a relevance score. This is the slowest stage per candidate, which is why it sees forty chunks and not four thousand, and it is the stage that turns "roughly related" into "actually answers this". Only the top handful after reranking reach the model. The reranker also produces the single most useful number in the pipeline: the score of the best chunk. That score is calibrated on our evaluation set, and it is the primary input to the refusal decision. ##### Citation grounding and the decision to refuse The generation prompt gives the model the reranked chunks, each labelled with an identifier, and asks for an answer in which every sentence that states a fact ends with the identifiers of the chunks that support it. Sentences that are not supported by any chunk are not allowed. The output is parsed, and each cited identifier is checked against the chunks actually provided; a citation to a chunk that was not in the context is treated as a hallucination and the answer is rejected. Grounding is then verified rather than trusted. For each cited sentence, a smaller model is asked whether the cited chunk supports that sentence, and the answer is only released if every sentence passes. When one sentence fails, the pipeline retries generation once with the failing sentence quoted back as a constraint; if it fails again the response falls back to a partial answer containing only the sentences that passed, clearly marked as partial. Refusal happens before generation when the evidence is weak, and after generation when the grounding check fails. The pre-generation gate looks at the best reranker score against the calibrated threshold and at the shape of the fused list: if the best chunk is below the threshold, or if the top candidates come from unrelated documents with no agreement, the system answers that it could not find this in the knowledge base and lists the closest documents it did find, so the user can judge whether the question was asked in the wrong words or the document genuinely does not exist. ```figure enterprise-rag-pipeline/grounding-gate The gate checks the best reranker score and verifies each cited claim against its span before an answer is released; otherwise the response is a refusal with the nearest documents. ``` ##### The evaluation harness came before prompt tuning We did not touch a prompt until we could measure retrieval and faithfulness. The evaluation set was built with the client: real questions from support logs and from staff, each paired with the chunk identifiers that contain the answer and a short reference answer. A second set held questions the knowledge base cannot answer, because a system that refuses correctly has to be scored on refusing. Two metrics were tracked per release. Retrieval recall at k asks whether the labelled chunks appear in the top k after fusion and after reranking; it tells us whether a wrong answer was the retriever's fault or the model's. Faithfulness asks a judge model whether every claim in the generated answer is supported by the retrieved chunks; it tells us whether the model stayed inside its evidence. Refusal accuracy on the unanswerable set is reported beside them, because it is trivial to improve faithfulness by refusing everything. ```python from statistics import mean def recall_at_k(retrieved_ids: list[str], relevant_ids: set[str], k: int) -> float: if not relevant_ids: return 1.0 hits = len(relevant_ids & set(retrieved_ids[:k])) return hits / len(relevant_ids) def evaluate(cases, retrieve, answer, judge, k: int = 5) -> dict[str, float]: recall, faithful, refusals_ok = [], [], [] for case in cases: chunks = retrieve(case["question"]) ids = [c.chunk_id for c in chunks] if case["relevant_ids"]: recall.append(recall_at_k(ids, set(case["relevant_ids"]), k)) result = answer(case["question"], chunks) if case["answerable"]: faithful.append(judge(result.text, [c.text for c in result.cited_chunks])) else: refusals_ok.append(1.0 if result.refused else 0.0) return { "recall_at_k": mean(recall) if recall else 0.0, "faithfulness": mean(faithful) if faithful else 0.0, "refusal_accuracy": mean(refusals_ok) if refusals_ok else 0.0, } ``` The harness runs in CI against a fixed index version and prints the three numbers next to the previous release's. A change that raises faithfulness by refusing more questions shows up immediately as a drop in recall or a change in refusal behaviour, and it gets discussed rather than merged. The judge itself needed checking. Early on it agreed with nearly everything, and when we read a sample of its verdicts by hand it was accepting sentences that paraphrased the chunk loosely enough to change the meaning. We rewrote the judge prompt to require the specific span that supports each claim and to answer no when it could not quote one, and we kept a small hand-labelled set of judge verdicts so that changes to the judge could themselves be scored. ##### Caching There are three caches and each one has a different key, which is the whole design. Embeddings are cached by content hash, so re-ingesting a corpus after a small edit only embeds the chunks that changed; this turned a full re-index from an evening job into minutes. Retrieval results are cached in Redis by normalised query text plus index version, with a short time to live, so a burst of the same question from a team reading the same announcement costs one retrieval. Reranker scores are cached by the pair of query hash and chunk hash, which matters because the same popular chunks appear in the candidate sets of many different queries. Nothing is cached across index versions. When the corpus changes, the index version increments, every retrieval key stops matching, and stale answers cannot survive by accident. ##### Observability Every request produces a trace with the query, the two raw ranked lists, the fused list, the reranker scores, the gate decision and its reasons, the prompt, the raw model output, the parsed citations, the grounding verdict per sentence, and the latency of each stage. Traces are written to PostgreSQL and are queryable, which is how most debugging happens: the question "why did it refuse this" is a query over the gate reasons, not a re-run. Aggregates are charted per release: refusal rate, share of refusals that were pre-generation versus grounding failures, reranker score distributions, and cache hit rates per cache. A sample of answered questions is pulled each week for human review, weighted towards low reranker scores that still passed the gate, because that is where a threshold that has drifted would show up first. ##### What broke along the way Prefixing the heading path to every chunk improved dense retrieval and quietly damaged keyword retrieval: the same heading words repeated in hundreds of chunks inflated their term frequencies and pushed genuinely relevant chunks down. The fix was to index the heading path as a separate weighted field rather than as part of the body. Changing the embedding model without re-indexing produced a week of nonsense, because query vectors and stored vectors were no longer in the same space. Index versioning, with the embedding model name as part of the version, exists because of that week. The refusal threshold was first set by eye and refused far too much. Calibrating it on the evaluation set, and reporting refusal accuracy beside faithfulness, is what made the threshold a decision rather than a mood. ##### What we would keep The evaluation harness before prompt work, a refusal path that is designed and measured rather than left to the model, chunking along document structure, and separate caches with separate keys. Reciprocal rank fusion earned its place by being the one part of the retrieval stack we never had to tune again. ### AdiNox: Building a TOTP Authenticator with a Biometric Gate - Stack: React, TypeScript, Supabase, face-api.js - Summary: How AdiNox parses otpauth QR codes, generates TOTP codes with WebCrypto, keeps secrets in an AES-GCM vault under a PBKDF2 key, and puts a face-api.js check in front of the code list without pretending a face is a key. - Challenge: An authenticator holds the one secret that can produce every future login code for an account, inside a browser where storage is readable by any script in the origin, screens are photographed, and devices get lost. AdiNox had to keep secrets encrypted at rest, keep codes off the screen when nobody is looking, survive device loss, and stay honest about what a webcam check can and cannot protect. - Solution: We parse otpauth URIs with the platform URL class and normalise every field, generate codes over crypto.subtle with non-extractable HMAC keys, and seal the vault with AES-GCM under a PBKDF2 key that lives only in memory. The face-api.js gate sits in front of the rendered codes within an unlocked session, cold starts always ask for the passphrase, the encrypted vault doubles as the backup format, and Supabase holds the account and preferences but never a secret. #### AdiNox: building a TOTP authenticator with a biometric gate AdiNox is a two-factor authentication app: it scans the QR code a service shows during enrolment, stores the shared secret on the device, and generates the six digit time-based codes you type back in. It is built in React and TypeScript, keeps its vault in the browser, uses Supabase for the account, and puts a face-api.js check in front of the code list. This is an account of the decisions behind it, in particular where the security actually lives and where it deliberately does not. ##### The threat model An authenticator is a small program with a large responsibility, so the first thing we wrote was not code but a list of what could go wrong. The asset is the set of TOTP secrets. Anyone holding a secret can produce every future code for that account, so a leaked secret is worse than a leaked code by a wide margin. Three attacks mattered to us. The first is the secret at rest: browser storage is readable by anything that runs in the same origin, is copied into profile backups, and sits on disk in the clear unless we do something about it. The second is the screen: a code that is visible is a code that can be photographed, screenshotted, or read over a shoulder, and a code list that stays open on an unattended laptop is an open door for thirty seconds at a time. The third is loss of the device itself, which combines the first two with the fact that the owner now needs a way back into their accounts. Two things were explicitly out of scope. TOTP is phishable: a user who types a valid code into a fake login page has handed it over, and no authenticator can prevent that. And a fully compromised browser, with a malicious extension reading the DOM while the vault is open, defeats any client-side design. We wrote both down so that later features would not pretend to solve them. ##### Scanning a QR code and parsing the otpauth URI Enrolment starts with the camera. The QR payload is an `otpauth://totp/` URI, and the format is loose enough that parsing it by regular expression is a mistake we made once and reverted. Labels are percent-encoded and may contain an issuer prefix separated by a colon, the `issuer` query parameter may disagree with that prefix, secrets arrive in upper or lower case base32 with or without padding, and some services include spaces. We parse with the platform `URL` class and normalise every field before it goes anywhere near storage. ```ts export interface TotpAccount { issuer: string; account: string; secret: Uint8Array; algorithm: 'SHA-1' | 'SHA-256' | 'SHA-512'; digits: 6 | 8; period: number; } const BASE32 = 'ABCDEFGHIJKLMNOPQRSTUVWXYZ234567'; export function decodeBase32(input: string): Uint8Array { const clean = input.toUpperCase().replace(/[=\s]/g, ''); const out: number[] = []; let bits = 0; let value = 0; for (const ch of clean) { const idx = BASE32.indexOf(ch); if (idx < 0) throw new Error('secret is not base32'); value = (value << 5) | idx; bits += 5; if (bits >= 8) { bits -= 8; out.push((value >>> bits) & 0xff); } } return Uint8Array.from(out); } export function parseOtpauth(uri: string): TotpAccount { const url = new URL(uri); if (url.protocol !== 'otpauth:' || url.host !== 'totp') throw new Error('only otpauth://totp is supported'); const label = decodeURIComponent(url.pathname.replace(/^\//, '')); const colon = label.indexOf(':'); const labelIssuer = colon >= 0 ? label.slice(0, colon).trim() : ''; const account = (colon >= 0 ? label.slice(colon + 1) : label).trim(); const q = url.searchParams; const secretParam = q.get('secret'); if (!secretParam) throw new Error('secret parameter missing'); const algoParam = (q.get('algorithm') ?? 'SHA1').toUpperCase().replace('SHA', 'SHA-'); const algorithm = algoParam === 'SHA-256' || algoParam === 'SHA-512' ? algoParam : 'SHA-1'; const digits = q.get('digits') === '8' ? 8 : 6; const period = Number(q.get('period') ?? 30); if (!Number.isInteger(period) || period < 15 || period > 120) throw new Error('unreasonable period'); return { issuer: (q.get('issuer') ?? labelIssuer).trim(), account, secret: decodeBase32(secretParam), algorithm, digits, period, }; } ``` The bounds check on `period` is there because a malformed QR with a period of zero would otherwise make the counter division blow up in a way that is confusing to debug from a user report. The decoded secret is a `Uint8Array`, never a string, from this point on. ```figure adinox-authenticator/secret-lifecycle The secret goes from the QR payload through the parser into an AES-GCM encrypted vault, and only leaves it as a non-extractable HMAC key that produces the six digit code. ``` ##### Generating codes with WebCrypto TOTP is HOTP with the counter set to the number of periods since the Unix epoch. HOTP is an HMAC over the eight byte big-endian counter, followed by dynamic truncation: take the low four bits of the last byte as an offset, read four bytes from that offset, mask the top bit, reduce modulo ten to the power of the digit count. All of this is a few lines over `crypto.subtle`, and we chose to write those lines rather than pull in a library, because the surface we most wanted to be able to read end to end was exactly this one. The important detail is how the key is imported. `crypto.subtle.importKey` is called with `extractable` set to false, so the HMAC key object that lives in memory during a session cannot be exported back to bytes by any script, including ours. The raw secret bytes are needed only twice: once when parsing the QR, and once when re-importing after the vault is decrypted. Both are short lived. ##### Encryption at rest and the key derivation choice The vault is a single JSON document of accounts, serialised and encrypted with AES-GCM under a key derived from the user's passphrase. GCM gives us authentication as well as confidentiality, so a tampered vault fails to decrypt instead of producing garbage accounts. Each write uses a fresh twelve byte random nonce, and the nonce is stored next to the ciphertext. The derivation function was the real decision. Argon2id is the better function against offline guessing, but WebCrypto does not ship it, so using it means a WebAssembly build, a larger bundle, and a memory cost we would have to tune per device. PBKDF2 with SHA-256 is native, runs off the main thread inside the browser's crypto implementation, and is well understood. We chose PBKDF2 with a high iteration count and a sixteen byte random salt, and accepted that the passphrase, not the KDF, is the main defence against an attacker who has copied the ciphertext. The iteration count is stored with the vault so it can be raised on the next passphrase change without a migration. ```ts const enc = new TextEncoder(); const dec = new TextDecoder(); export interface SealedVault { v: 1; kdf: { name: 'PBKDF2'; hash: 'SHA-256'; iterations: number; salt: string }; iv: string; data: string; } const b64 = (b: Uint8Array) => btoa(String.fromCharCode(...b)); const unb64 = (s: string) => Uint8Array.from(atob(s), (c) => c.charCodeAt(0)); export async function deriveVaultKey(passphrase: string, salt: Uint8Array, iterations: number): Promise { const material = await crypto.subtle.importKey('raw', enc.encode(passphrase), 'PBKDF2', false, ['deriveKey']); return crypto.subtle.deriveKey( { name: 'PBKDF2', hash: 'SHA-256', salt, iterations }, material, { name: 'AES-GCM', length: 256 }, false, ['encrypt', 'decrypt'], ); } export async function sealVault(key: CryptoKey, accounts: unknown, kdf: SealedVault['kdf']): Promise { const iv = crypto.getRandomValues(new Uint8Array(12)); const plain = enc.encode(JSON.stringify(accounts)); const cipher = await crypto.subtle.encrypt({ name: 'AES-GCM', iv }, key, plain); return { v: 1, kdf, iv: b64(iv), data: b64(new Uint8Array(cipher)) }; } export async function openVault(key: CryptoKey, sealed: SealedVault): Promise { const plain = await crypto.subtle.decrypt({ name: 'AES-GCM', iv: unb64(sealed.iv) }, key, unb64(sealed.data)); return JSON.parse(dec.decode(plain)) as T; } ``` The derived key is also non-extractable. It exists in memory while the tab is open and unlocked, and it is dropped on explicit lock, on a period of inactivity, and when the page unloads. A cold start always asks for the passphrase. We considered caching the key in `sessionStorage` to survive a reload and rejected it, because anything we can read back as bytes an injected script can read too. Storage itself is IndexedDB rather than `localStorage`. The reasons are practical: it handles binary without base64 inflation, it is asynchronous so a large vault does not block the frame, and it is less likely to be trimmed under storage pressure. ##### The face-api.js gate, and what it is not The biometric check is the feature people ask about first, so it is worth being precise about it. face-api.js runs entirely in the browser: a detector finds a face in the webcam frame, a landmark model aligns it, and a recognition model produces a 128 dimensional descriptor. At enrolment we store a small set of descriptors; at unlock we compute a fresh one and compare it by Euclidean distance against the enrolled set under a threshold. That descriptor is not a key. It varies from frame to frame, so it cannot be fed to a KDF, and even if it could, a face is not a secret. The gate therefore sits in front of the code list, not in front of decryption. Within an unlocked session, when the inactivity timer fires or the user taps lock, the vault key stays in memory but the interface hides every code behind the camera check. This gives the thing people actually want, which is not having to type a long passphrase every time they pick up the laptop, without pretending the face is protecting the ciphertext. What it protects against is the second item on the threat list: a colleague glancing at an unattended screen, a code list left open in a shared space. What it does not protect against is an attacker with the device and a photograph, a compromised page, or anyone with the vault file, and the settings screen says so in plain words. The threshold is a tradeoff between false rejections in poor light and false acceptances of similar faces; we set it conservatively and always leave the passphrase as the fallback, since a gate that locks out its owner is a gate that gets disabled. ```figure adinox-authenticator/unlock-flow A cold start goes through the passphrase and PBKDF2 to a non-extractable key; within a session the face-api.js gate only stands between the user and the rendered codes. ``` The models are loaded lazily, only when a user enables the gate, because the weights are large relative to the rest of the app and most of the time the camera is not needed. ##### Clock drift A TOTP code is a function of the shared secret and of the current thirty second window, so the device clock is part of the protocol. Servers typically accept the current window and one on either side, which gives about a minute of tolerance in total; a device that is two minutes fast produces codes that are simply wrong, and the failure is silent. The app does three things about this. It computes the counter from `Date.now()` and re-renders exactly on the window boundary rather than on a fixed tick, so the countdown ring and the code are never out of step with each other. It fades the code in the last few seconds of a window and shows the next one immediately, because a code copied with two seconds left is likely to be rejected by the time it is submitted. And on unlock it compares the local clock against the `Date` header of a request to its own origin, and shows a warning when the skew exceeds a few seconds. The warning does not correct the clock, since generating codes against a time the user cannot see would be more confusing than the drift; it tells them what to fix. ```figure adinox-authenticator/code-rotation Codes rotate on thirty second window boundaries; a validating server accepts the neighbouring windows, so small drift is tolerated and large drift fails silently. ``` ##### Backup and export Device loss is the case that decides whether an authenticator is trusted. The backup format is the sealed vault itself: the same JSON with salt, iterations, nonce and ciphertext, downloadable as a file and importable on another device with the same passphrase. Nothing in it is readable without the passphrase, so a backup left in a downloads folder is no worse than the ciphertext already on disk. A plaintext export exists for migrating to another authenticator, one account at a time, as the original `otpauth` URI rendered as a QR code on screen. It is behind the passphrase, not the face gate, and it is per account rather than bulk, because a bulk plaintext export is the kind of file that ends up in a chat history. ##### The role of Supabase Supabase holds the account: sign-in, profile, and preferences such as whether the biometric gate is enabled and the inactivity timeout. It does not hold secrets, and it does not hold the vault. We did design a sync of the sealed vault, since the blob is opaque to the server, and did not ship it in this version. The reason is the key: a vault synced across devices is only as safe as the passphrase on the weakest of them, and the moment the ciphertext is on a server, an attacker has unlimited time to guess against PBKDF2. Until the derivation is stronger, the vault stays where it was created, and the export file is the way to move it. That division also keeps the recovery story honest. Losing the Supabase password loses preferences. Losing the vault passphrase loses the codes, and the app says so at enrolment, because the alternative, a server that can recover your secrets, would mean the server could read them. ### Farm Digitisation and Parcel Mapping at District Scale - Stack: PostGIS, GeoPandas, GDAL, QGIS, Python - Summary: Four disagreeing descriptions of the same fields, GPS tracks, scanned cadastral sheets, operator polygons and model boundaries, were georeferenced, conflated with confidence scoring and reviewed by people, then delivered as a versioned parcel layer to the client's GIS and to PostGIS. - Challenge: Zaraat Dost Private Limited needed one authoritative parcel layer for a district, but the parcels already existed in four incompatible forms: partial GPS survey tracks, scanned paper cadastral sheets with no coordinates, quickly drawn operator polygons in Web Mercator, and dense HRNet-derived boundaries that knew nothing about ownership. None was complete, none was fully trustworthy, and every downstream system would store whatever identifiers we issued. - Solution: We reprojected everything into one metric CRS, warped the paper sheets with ground control points and a thin plate spline, and treated every input polygon as a scored candidate rather than as truth. A PostGIS conflation query clustered overlapping candidates by intersection over union, a confidence function picked a winner and kept alternates, and close calls went to a QGIS review queue where every decision wrote a new parcel version. Exports to GeoPackage, Shapefile and PostGIS carried stable identifiers with versions, and a field audit fed back into the source priors. #### Farm digitisation and parcel mapping at district scale Zaraat Dost Private Limited needed one authoritative parcel layer for an entire district: every farm as a polygon, every polygon tied to a farmer record, and every edit traceable. The difficulty was never drawing polygons. It was that four different descriptions of the same fields already existed, they disagreed with each other, and none of them was complete on its own. This case study is about the pipeline that took those sources in, reconciled them, put the arguments in front of a human, and delivered a versioned layer to the client's GIS and to a PostGIS store. ##### Four sources that describe the same ground differently The intake was heterogeneous in every sense that matters to a geospatial engineer. Field survey GPS tracks came from handheld GNSS units carried by enumerators walking parcel boundaries. They were the most trusted source for the parcels they covered, but coverage was partial, the tracks contained loops where the enumerator retraced steps, and a few units had been configured to log in a local datum rather than WGS84, which we only discovered when a batch of tracks sat a consistent distance off the basemap. Scanned cadastral sheets were paper maps photographed or flatbed scanned. They carried no coordinate reference at all, had different scales from sheet to sheet, and showed boundaries as they were recorded on paper years earlier. They were valuable precisely because they encoded legal parcel divisions that no imagery model would ever recover, such as a field split between two heirs with no visible fence. Operator-drawn polygons had been digitised by the client's team in a web map over a basemap. That put them in Web Mercator, with whatever registration error the basemap tiles carried in that area. Some were careful; some were rectangles drawn quickly to get a record into the system. Model-derived boundaries came from our HRNet-W48 field boundary delineation work, vectorised from the boundary raster into GeoJSON. They were dense and complete, but the model has no idea about ownership: two terraces farmed by the same family often came out as one polygon, and a single legal parcel with a bund through the middle came out as two. Each source was strong where the others were weak, so the decision that shaped everything else was that no source would be treated as ground truth. Each candidate polygon would carry a confidence, and the pipeline would choose. ##### Georeferencing and one working CRS Before any geometry could be compared, all four sources had to land in the same coordinate reference system, and it had to be a metric one. Area, overlap ratios and buffer distances are meaningless in degrees, and several of our quality checks are thresholds in metres. We settled on UTM zone 43N (EPSG:32643) as the working CRS, with WGS84 (EPSG:4326) used only at the export boundary because the client's GIS expected it. GPS tracks were parsed from GPX and CSV with GeoPandas, assigned EPSG:4326, then reprojected. The datum problem was caught by a sanity check rather than by cleverness: for every batch we computed the median offset between track vertices and the nearest model boundary, and any batch with an offset above a few metres was quarantined for inspection instead of being conflated. Operator polygons came in as EPSG:3857 and were reprojected directly. The basemap registration error could not be corrected globally, so it was folded into the source's confidence prior rather than fixed. The cadastral sheets were the real georeferencing work. For each sheet an operator placed ground control points at features that existed both on paper and in imagery, typically road junctions, canal crossings and survey stones. We then warped the scan with GDAL. A first-order polynomial (an affine transform) was tried first because it is well behaved, but the paper had been folded and the scans were not flat, so residuals at the sheet edges were unacceptable. A thin plate spline warp absorbed the local distortion and brought the residuals down to a level that matched the line weight on the paper itself. ```bash gdal_translate -of GTiff \ -gcp 412 388 74.3121 31.5210 \ -gcp 3810 402 74.3407 31.5203 \ -gcp 3792 2650 74.3401 31.4998 \ -gcp 420 2668 74.3118 31.5004 \ sheet_17.tif sheet_17_gcp.tif gdalwarp -tps -r bilinear -t_srs EPSG:32643 \ -co COMPRESS=DEFLATE sheet_17_gcp.tif sheet_17_utm.tif ``` The warped sheet was then traced into polygons by an operator in QGIS, and each polygon inherited the sheet's residual statistics as part of its confidence. A sheet with poor residuals produced candidates that could still win, but only when nothing better existed. ```figure farm-digitisation-pipeline/crs-alignment A scanned sheet is pinned to ground control points and warped, then every source is reprojected into the metric working CRS before comparison. ``` ##### Conflation: finding the arguments before settling them With everything in EPSG:32643 and loaded into PostGIS, conflation became a question of which candidate polygons were describing the same parcel. We used intersection over union rather than plain intersection area because IoU is symmetric and scale-free: a large model polygon that happens to swallow a small survey polygon gets a low IoU, which is exactly right, whereas raw intersection area would have flattered it. The query below produces candidate pairs across different sources. It runs on a spatially indexed table and only computes the expensive intersection when the bounding boxes overlap. ```sql with candidates as ( select a.candidate_id as id_a, b.candidate_id as id_b, a.source as source_a, b.source as source_b, st_area(st_intersection(a.geom, b.geom)) / nullif(st_area(st_union(a.geom, b.geom)), 0) as iou from parcel_candidates a join parcel_candidates b on a.candidate_id < b.candidate_id and a.source <> b.source and a.geom && b.geom and st_intersects(a.geom, b.geom) where st_isvalid(a.geom) and st_isvalid(b.geom) ) select * from candidates where iou >= 0.35 order by iou desc; ``` Two details in that query came from failures. The `st_isvalid` filter exists because a self-intersecting operator polygon crashed a nightly run with a topology exception inside `st_intersection`; we now run `st_makevalid` at ingest and reject anything that is still invalid. The `nullif` guard exists because a pair of degenerate slivers produced a zero-area union and a division error. Pairs were grouped into clusters with a union-find pass in Python, so that a survey polygon, an operator polygon and a model polygon that all overlapped each other became one cluster with three candidates. A cluster then needed a single winner. We tried averaging boundaries and abandoned it within a day: averaging a terraced parcel's survey boundary with a model boundary that had merged two terraces produces a shape that matches neither reality nor any record. Instead, the pipeline scores each candidate and picks the best one, keeping the others as alternates so a reviewer can switch. ```python from dataclasses import dataclass SOURCE_PRIOR = {"survey": 0.90, "cadastral": 0.70, "operator": 0.60, "model": 0.55} @dataclass class Candidate: candidate_id: str source: str area_m2: float sliver_ratio: float # perimeter**2 / area, high for thin fragments georef_rmse_m: float # 0 for sources with no warp step support_iou: float # best IoU with a candidate from another source def confidence(c: Candidate) -> float: prior = SOURCE_PRIOR[c.source] shape = 1.0 / (1.0 + max(0.0, c.sliver_ratio / 80.0)) georef = 1.0 / (1.0 + c.georef_rmse_m / 3.0) support = 0.5 + 0.5 * min(c.support_iou, 1.0) return prior * shape * georef * support def merge_cluster(cluster: list[Candidate], review_margin: float = 0.08): ranked = sorted(cluster, key=confidence, reverse=True) best, rest = ranked[0], ranked[1:] needs_review = bool(rest) and confidence(best) < confidence(rest[0]) + review_margin return best, rest, needs_review ``` The `support` term is what makes conflation more than a lookup of source priors. A model polygon that agrees closely with an operator polygon is stronger evidence than either alone, and a survey track that nothing else agrees with is worth a second look even though survey is the most trusted source. The `review_margin` is the honest part: when the top two candidates are within a small margin, the pipeline does not pretend to know, it routes the cluster to a human. ```figure farm-digitisation-pipeline/conflation Candidates from four sources are clustered by intersection over union, scored, and merged into one parcel layer where each polygon carries a confidence. ``` ##### Linking parcels to farmer records The client held a farmer registry with a record identifier, a village code and, for surveyed farms, the survey job number. Linking a merged parcel to a farmer used the strongest key available and fell through to weaker ones. A parcel whose winning candidate came from a survey track carried the job number and linked directly. Parcels without that link were spatially joined to survey points and to cadastral sheet references, both of which named a farmer identifier. The remaining parcels were matched by village code plus a normalised name comparison against the registry, and every match from that path was marked as provisional. One-to-many results were common and legitimate: a farmer with three separate fields, or a field jointly farmed by two households. The pipeline never collapsed these. It stored a link table with a link method and a link confidence, and any parcel with more than one high-confidence farmer, or none at all, went to the review queue alongside the geometric conflicts. ##### The review queue The review queue was a PostGIS table of open items, each with a parcel identifier, a reason code and the alternates, published as a layer in a QGIS project. Reasons included close confidence margins, a survey candidate with no support from any other source, invalid or sliver geometry that survived cleaning, and ambiguous farmer links. Reviewers saw the winning geometry and its alternates on top of the imagery, and their tools were deliberately limited: accept, switch to an alternate, redraw, or reject with a reason. Every decision was written back as a row with the reviewer identity, the action, the reason and a timestamp, and the decision itself created a new parcel version rather than overwriting anything. This mattered in practice. Two reviewers in different weeks reached different conclusions about the same terraced slope; because both decisions were recorded with reasons, the disagreement was visible and could be settled by a field visit instead of by whoever edited last. ```figure farm-digitisation-pipeline/review-versioning Conflicts flow from the queue to a reviewer, each decision writes a new version to the ledger, and audit samples feed back into the queue. ``` ##### Versioned parcel identifiers Parcel identifiers had to survive edits, splits and merges without ever being reused, because the client's downstream systems would store them. Each parcel received a stable identifier at the moment its cluster was first merged, and a monotonically increasing version number. A redraw produced a new version of the same identifier. A split retired the parent version and created child parcels with new identifiers whose lineage rows pointed at the parent. A merge did the reverse. The current layer was simply a view over the version table filtered to the latest non-retired version of each identifier, so history was never a separate system to maintain. Exports carried the identifier and the version together, and the client's GIS was told to treat the pair as the key. When a later delivery changed a boundary, the client could see which version they held and what had replaced it. ##### Exports to the client's GIS and to PostGIS The PostGIS store was the source of truth: parcels with a `geometry(MultiPolygon, 32643)` column, a GiST index, check constraints on validity, and the version and lineage tables beside it. The client's GIS received GeoPackage in EPSG:4326 as the primary deliverable, with Shapefile alongside because some of their tooling still expected it. Shapefile cost us a debugging afternoon: its ten-character field name limit silently truncated `farmer_link_confidence` and `georef_rmse_m` into colliding names until we added an explicit field mapping and a test that round-tripped a sample export and compared attributes. ##### The QA loop Automated checks ran on every build of the layer: geometry validity, overlaps between current parcels above a small tolerance, gaps along shared boundaries, area outside a plausible range for the district, duplicate identifiers, and any parcel whose winner had changed source between versions without a review decision attached. Anything flagged went into the same review queue, so reviewers had one place to work. The second half of the loop was a field audit. A random sample of parcels, weighted towards those decided by the weakest evidence, was visited and walked with a GNSS unit. Where the walked boundary disagreed with the delivered one, the parcel was corrected as a new version and the disagreement was recorded against the source that had won. Those records were how the source priors in the confidence function were adjusted over time. The priors started as our judgement and ended as something closer to a measurement of each source's actual reliability in that district. ##### What we would keep and what we would change Treating every input as a scored candidate rather than as truth was the decision that made the rest tractable, and the review margin that admits uncertainty is what kept reviewer time focused on real disagreements. The thin plate spline warp was the correct choice for folded paper. Versioning from the first merge, rather than adding it later, is the one thing we would insist on again without discussion. We would start the field audit earlier. The priors were guesses for longer than they needed to be, and a small audit sample in the first weeks would have corrected them sooner and reduced the number of clusters that reached the review queue. ### AdiGaze: AI Recruitment Engine - Stack: React (Vite), Supabase (Deno), Gemini 1.5 Pro, pgvector - Summary: AdiGaze turns a batch of mixed-format resumes into structured, ranked candidate records, and it was built to keep finishing the batch when a provider throttles or a file takes minutes to parse. - Challenge: High-volume resume parsing runs into two hard limits at once: serverless functions time out long before a large batch completes, and hosted LLM APIs return 429s under sustained load. We needed throughput and resilience without a dedicated queue service or a bigger budget. - Solution: We spread work across whatever API keys are configured using a bounded round-robin worker pool inside Supabase Edge Functions, streamed results over Server-Sent Events so the shortlist appears as each candidate lands, and added exponential backoff with a keyword-search fallback. A SQL pre-filter narrows the set before pgvector ranks it. #### AdiGaze: an AI recruitment engine that keeps parsing when the API says no AdiGaze takes a pile of resumes in mixed formats and turns them into structured candidate records, then ranks those candidates against a role. The interesting engineering was never the parsing prompt. It was making a large batch finish reliably on a serverless budget, when the two constraints that matter most (function timeouts and provider rate limits) pull in opposite directions. ##### Who it was for The product is aimed at small recruiting teams who upload dozens or hundreds of resumes at once and expect a ranked shortlist in the same sitting. That sets the bar: a recruiter should see results appear continuously, not watch a spinner for minutes and then receive everything at once (or, worse, a timeout). ##### The constraints we designed around - **Serverless timeouts.** Supabase Edge Functions have a short execution budget. A single request cannot sit and parse a hundred PDFs. - **Provider rate limits.** Hosted LLMs cap requests per minute. Fire a whole batch at once and you collect `429`s. - **Budget and team size.** No dedicated queue service, no ops team. Whatever we built had to live inside Supabase and stay maintainable by one person. ##### Architecture ```mermaid flowchart LR U[Recruiter uploads batch] --> ST[(Supabase Storage)] U --> EF[Edge Function: parse] EF --> Q{Work queue} Q --> W1[Worker lane 1] Q --> W2[Worker lane 2] Q --> W3[Worker lane N] W1 --> G[LLM provider] W2 --> G W3 --> G W1 --> DB[(Postgres + pgvector)] EF -- SSE --> U DB --> RANK[Pre-filter then vector rank] RANK --> U ``` ##### The core decision: a bounded worker pool instead of a queue service The obvious answer to "process many things without timing out" is a managed queue. I rejected that early. It meant new infrastructure, another bill, and another moving part to secure, all for batch sizes that rarely reach the thousands. Instead I kept the whole thing inside the Edge Function and treated the set of configured API keys as the unit of parallelism. The function discovers how many keys are present at runtime and opens a fixed number of lanes per key. Each lane pulls from a shared in-memory queue until the batch is drained. The point is not to maximize raw fan-out; it is to keep concurrency at a level the provider will actually tolerate. ```ts // Discover keys at runtime so scaling means adding an env var, not a deploy. const keys = [1, 2, 3, 4, 5] .map((i) => Deno.env.get(`GEMINI_API_KEY_${i}`)) .filter((k): k is string => Boolean(k)); if (keys.length === 0) throw new Error("No AI provider keys configured"); const LANES_PER_KEY = 2; const queue = new BatchQueue(files, { size: 8 }); await Promise.all( keys.flatMap((key) => Array.from({ length: LANES_PER_KEY }, () => drain(queue, key)), ), ); ``` Here are the options I weighed: | Option | Why not / why yes | | --- | --- | | External queue (SQS, Cloud Tasks) | New infra and cost; overkill for these batch sizes | | One request per file, sequential | Simple but far too slow; the recruiter waits minutes | | Fan everything out with `Promise.all` | Immediate `429` storms; no backpressure | | Bounded pool across available keys | Backpressure and parallelism with zero extra infrastructure (chosen) | ##### Streaming results with Server-Sent Events A plain REST endpoint waits for the whole job before it answers. I switched the parse endpoint to Server-Sent Events so each finished candidate is pushed the moment it is ready. That changes what the recruiter feels: perceived latency collapses from "the entire batch" to "the first candidate." It is also a reliability feature: a long running connection that emits progress is far less likely to be killed as an idle request than one that goes silent for minutes. ##### Ranking: cheap filter first, vectors second Semantic search over every stored candidate is wasteful when most of them fail a hard requirement anyway. So ranking runs in two stages: a SQL prefilter on non-negotiable criteria, then a vector similarity search on whatever survives. ```sql select id, full_name, 1 - (embedding <=> $1) as score from candidates where user_id = auth.uid() and location = $2 and work_authorized = true order by embedding <=> $1 limit 25; ``` Filtering first keeps the vector scan small and the query fast. I measured query latency before and after adding the prefilter rather than trusting intuition, and the pre-filtered path was consistently the cheaper one. ##### The hardest tradeoff: partial failure In a batch of a hundred, something will fail: a malformed PDF, a transient `429`, a provider hiccup. The design rule was that one bad item must never sink the batch. A lane that hits a rate limit backs off with exponential delay and returns its item to the queue for another lane to retry. If the model is unreachable after retries, the item falls back to a keyword extraction path and is flagged so a human can review it later. The system degrades to "less clever" rather than to "broken." ##### How it was tested - Synthetic batches across a range of sizes, including deliberately corrupt files. - A throttled key to confirm that backoff and re-queue actually fire under `429`s. - A forced worker error mid-stream to confirm the SSE connection and the rest of the batch survive. - Row Level Security policy tests so a user can only ever read candidates tied to their own `user_id`. - Input sanitization that strips control and zero-width characters before text reaches either the model or the database. ##### What shipped Batch parsing that finishes reliably at the sizes real users upload, a shortlist that streams in live, and graceful degradation to keyword search when the model is unavailable. ##### Lessons 1. Size concurrency to your rate limit, not your CPU count. 2. Streaming is a reliability feature, not only a UX nicety. 3. A cheap prefilter is what makes vector search affordable. 4. Design the failure path before the happy path. 5. Sanitize untrusted input before the model sees it, not after. ### The Agentic Blueprint - Stack: System Design, AI, Swarm Logic, pgvector, Resilience - Summary: A field guide to building AI systems that actually finish work: five patterns (structured cognition, parallel workers, self-healing reflexes, vector memory, and an input immune system) drawn from the recruitment engine, RAG pipeline, and MCP server. - Challenge: Most teams treat a language model as a vending machine (prompt in, text out) and then struggle to make it act reliably, run concurrently, or recover from failure. The challenge is turning a passive model into a system that decides, uses tools, and heals itself. - Solution: A five-part blueprint: constrain the model to a schema so its output is data, run a bounded pool of workers sized to the rate limit, add backoff and graceful fallback, give it vector memory paired with a cheap SQL pre-filter, and sanitize every input while the database enforces least privilege. #### The agentic blueprint: turning an LLM into a system that finishes work Most people use a language model as a vending machine: prompt in, text out. An agent is different. It decides what to do next, it uses tools, it recovers from failure, and it runs more than one thing at a time. This is the pattern I keep returning to across the recruitment engine (AdiGaze), the RAG pipeline, and the MCP tool server: five parts that turn a model into something that actually completes a job. ##### The shape of an agent ```mermaid flowchart TD IN[Untrusted input] --> SAN[Sanitize and validate] SAN --> COG[Cognition: structured decision] COG --> TOOL{Call a tool?} TOOL -- yes --> ACT[Tool call with schema check] ACT --> OBS[Observe result] OBS --> COG TOOL -- no --> OUT[Return grounded answer] COG --> MEM[(Vector memory)] MEM --> COG ``` ##### Part 1: structured cognition The first failure mode of any agent is a model that rambles. If the output is prose, your code cannot act on it. So I constrain the model to a schema and treat its answer as data, not conversation. ```ts const CandidateJudgment = z.object({ matchScore: z.number().min(0).max(100), reasons: z.array(z.string()).max(3), }); ``` Validating against a schema like this does two things. It makes downstream logic deterministic (I can filter on `matchScore > 80` with confidence) and it caps token usage, because the model is not free to write an essay where a number and a short list will do. When validation fails, that is a signal to retry or fall back, not to crash. ##### Part 2: parallel workers A single agent is slow because it waits on the network for most of its life. The fix is to run several, each on its own credential, sharing one work queue. The number of workers should track what your provider will tolerate, not how many cores you have; the bottleneck is almost always the rate limit, not local compute. The queue gives you backpressure for free: workers pull the next item only when they finish the last, so you never fire a thousand requests at once. Adding capacity means adding a key and a lane, not rewriting the loop. Keeping the pool bounded also protects you from yourself. An unbounded fan-out will happily exhaust memory, trip the provider's concurrency ceiling, and turn one slow item into a cascade of timeouts. A fixed number of lanes per credential is a ceiling you choose deliberately, tuned to the quota you actually have rather than the one you wish you had. ##### Part 3: self-healing reflexes Agents run in a world where APIs time out and quotas run dry. A durable agent has reflexes: automatic responses to failure that need no human. - **Backoff and retry.** On a transient error, wait a little, then a little longer, then retry. This gives the upstream service room to recover instead of hammering it. - **Graceful fallback.** If the model is still unreachable after retries, do not error out. Switch to a simpler deterministic method, return a usable result, and mark it so a human can review it later. The goal is that the system degrades in quality rather than falling over. A recruitment tool that quietly drops to keyword matching during a provider outage is still doing its job. ##### Part 4: memory that means something A chatbot forgets everything when the tab closes. An agent needs memory that survives, and it should store meaning, not just strings. When a document arrives, I embed it and keep the vector alongside the row. Retrieval then works by similarity: a search for "frontend engineer" surfaces a candidate who lists React and TypeScript even if that exact phrase never appears. The trap here is retrieving too much. Vector search is only cheap when the candidate set is already small, so I pair it with a SQL prefilter on hard constraints. Filter to what is possible, then rank by meaning. ##### Part 5: the immune system Agents ingest untrusted data (uploaded files, form fields, external submissions), and that data flows straight toward a model and a database. Two defenses are non-negotiable: - **Sanitize before the model.** Strip control and zero-width characters, coerce types (`coerceInt`, `sanitizeStringArray`), and reject anything that does not match the expected shape. Treat instructions found inside ingested content as data, never as commands. - **Least privilege at the data layer.** Row Level Security so a tool call can only ever touch rows the current user owns. A confused or manipulated agent should still be boxed in by the database. ##### Where this maps to real systems | Part | In practice | | --- | --- | | Cognition | Zod/JSON schemas on every model response | | Workers | Bounded pool across API keys inside an Edge Function | | Reflexes | Exponential backoff plus a keyword fallback path | | Memory | pgvector embeddings with a SQL prefilter | | Immune system | Input sanitization and Row Level Security | ##### The tradeoffs worth naming Structure costs flexibility: a strict schema occasionally rejects a valid but oddly phrased answer, so the retry path has to be forgiving. Parallelism costs determinism: results arrive out of order, which is why streaming and stable sorting matter. Memory costs storage and staleness: embeddings drift as the underlying model changes, so re-embedding on a schedule has to be part of the plan. None of these are reasons to avoid the pattern; they are the things to budget for. ##### Lessons 1. Force structure early: an agent you cannot parse is an agent you cannot trust. 2. Concurrency is a rate-limit problem, not a compute problem. 3. Plan the failure path first; it is most of the engineering. 4. Store meaning, but retrieve narrowly. 5. Treat every ingested byte as untrusted, and let the database enforce the boundary. ### The Supabase Sovereign - Stack: Supabase, Postgres, System Design, Resilience, Edge Functions - Summary: How I keep Postgres-backed backends fast and secure on Supabase under serverless load: connection pooling, targeted indexing, edge compute, async webhooks, and Row Level Security, with the tradeoff behind each decision. - Challenge: Serverless frontends spawn thousands of short-lived functions, each wanting a database connection, and a modestly sized Postgres hits its connection ceiling long before it runs out of CPU. Left unmanaged, that pattern takes the whole backend down under a traffic spike. - Solution: Put a transaction-mode pooler in front of the database, index for the queries actually run, keep heavy logic in Edge Functions off the shared database CPU, push slow side effects to async pg_net webhooks, and treat default-deny Row Level Security as the real authorization boundary. #### Building resilient backends on Supabase Across the products I ship on Supabase (the AdiCorp HRMS, the Federals news CMS, the Aditron chat app, the AdiNox authenticator), the hard part was never storing data. It was surviving the connection pattern that serverless frontends produce: thousands of short lived functions, each wanting a database connection, all at once. This is a walk through the decisions that keep a Postgres-backed backend fast and secure under that load, and the tradeoffs behind each one. ##### The request path ```mermaid flowchart LR C[Client] --> API[PostgREST / Edge Function] API --> POOL[Supavisor pooler] POOL --> PG[(Postgres)] PG --> RLS[Row Level Security] PG --> IDX[Indexes] API --> RT[Realtime channel] PG -- pg_net --> HOOK[Async webhook] ``` ##### The connection wall, and how pooling gets past it Each direct Postgres connection costs real memory. In a serverless world the frontend scales horizontally, so a traffic spike turns into a spike of connection attempts, and a modestly sized database hits its connection ceiling long before it runs out of CPU. The fix is a pooler in front of the database. Supabase provides Supavisor; pointing the app at the pooler instead of the database directly changes the model from one connection per client to many clients sharing a small pool. The choice that matters is the mode: - **Session mode** holds a connection for the life of a client session. Use it when you need session-level features like prepared statements or `SET` state. - **Transaction mode** returns the connection to the pool the instant a statement finishes. This is what lets a handful of real connections serve a large number of concurrent, short lived function invocations. For serverless functions, transaction mode is almost always the right default. The tradeoff is that anything relying on session state has to be rethought, which is a feature, not a bug: it forces stateless request handling. ##### Indexing for the queries you actually run As tables like attendance logs in AdiCorp grow, an unindexed query degrades linearly until it dominates every page load. A few habits keep latency flat: - **Partial indexes** for the rows you query most. If the app almost always reads active records, index only those and keep the index small. - **The right index type for the shape of the data:** GIN for `jsonb` and full-text search, standard B-tree for equality and range lookups on ordinary columns. ```sql create index idx_attendance_active on attendance (employee_id, work_date) where status = 'active' and deleted_at is null; ``` I do not index speculatively. An unused index still has to be maintained on every write, so I add one in response to a slow query I have actually seen in the logs, then confirm the plan changed. ##### Thin database, fat edge Postgres will happily run business logic in PL/pgSQL, but heavy logic there consumes the one resource the whole system shares: database CPU. So I keep computation-heavy work (payroll runs, report generation) in Edge Functions written in Deno, and reserve database functions for logic that genuinely belongs next to the data. ```ts import { createClient } from "https://esm.sh/@supabase/supabase-js@2"; Deno.serve(async (req) => { const admin = createClient( Deno.env.get("SUPABASE_URL")!, Deno.env.get("SUPABASE_SERVICE_ROLE_KEY")!, ); const { employeeId, daysPresent, dailyRate } = await req.json(); const amount = daysPresent * dailyRate; const { error } = await admin .from("salaries") .upsert({ employee_id: employeeId, amount, updated_at: new Date() }); return new Response(JSON.stringify({ amount, ok: !error }), { status: error ? 400 : 200, headers: { "content-type": "application/json" }, }); }); ``` Secrets for these functions live in the Supabase vault, never in the database rows themselves, so SQL-level access does not hand an attacker the integration keys. ##### Doing slow work without making the user wait Synchronous triggers block the request that fired them. When a write should kick off something slow (notifying an external service, enriching a record), I use `pg_net` to fire an asynchronous HTTP request from the database and let the user's transaction return immediately. ```sql select net.http_post( url := 'https://example.internal/process', body := jsonb_build_object('id', new.id) ); ``` The tradeoff is that the work is now fire-and-forget, so it needs its own logging and retry story; an async webhook that silently fails is worse than a slow synchronous one. ##### Realtime, with the right primitive Supabase Realtime, built on Elixir, offers three tools, and picking the wrong one is a common mistake: - **Postgres changes** stream row changes to clients, good for keeping UI in sync with committed state. - **Broadcast** is ephemeral fire-and-forget messaging (a "user is typing" indicator in Aditron) that never touches the database, so it stays cheap. - **Presence** tracks who is online without race conditions. Routing a typing indicator through database changes would generate write load for information nobody needs to persist. Broadcast is the correct, cheaper choice. ##### Security as a default, not a layer Every table starts with Row Level Security on and a default-deny posture. In AdiCorp an employee can read only their own salary row while an admin sees the aggregate; in Federals the editorial tables are locked to authenticated roles. Because PostgREST exposes the schema directly to the client, RLS is not an optimization: it is the entire authorization boundary, and it has to be right. ```sql create policy "own salary only" on salaries for select to authenticated using (employee_id = auth.uid()); ``` ##### Comparing the choices | Decision | Alternative rejected | Why | | --- | --- | --- | | Transaction-mode pooling | Direct connections | Direct connections exhaust memory under serverless spikes | | Edge Functions for heavy logic | PL/pgSQL everywhere | Keeps shared database CPU free | | Async `pg_net` webhooks | Synchronous triggers | Users should not wait on downstream calls | | RLS default-deny | App-layer checks only | PostgREST exposes the schema; the database must enforce access | ##### Lessons 1. Pool connections in transaction mode before you scale, not after the first outage. 2. Add indexes in response to real slow queries, and confirm the plan changed. 3. Keep heavy computation off the database's shared CPU. 4. Make slow side effects asynchronous, but give them their own observability. 5. Treat Row Level Security as the authorization boundary, because with PostgREST it is. ## Notes Security and engineering write-ups published at https://adilmunawar.vercel.app/#blog. ### What Actually Makes LLM Inference Fast: KV Cache, Batching and Speculative Decoding - Tags: LLM, Inference, GPU, Systems - Summary: Why decode is bound by memory bandwidth rather than compute, how large the KV cache really gets, what paged attention and continuous batching recover, how speculative decoding speeds up generation without changing the output distribution, and the three numbers to measure before believing any serving benchmark. #### What Actually Makes LLM Inference Fast: KV Cache, Batching and Speculative Decoding Most of the running cost of the systems we build now, from the retrieval pipeline that answers questions over a company's documents to the agent loop that calls tools through an MCP server, is spent inside a language model producing tokens one at a time. Training gets the papers; inference gets the bill. When we first paid serious attention to serving cost, our mental model was that a faster GPU means faster generation and that a model which fits in memory is served well. Both were wrong in instructive ways, and this note is the explanation we wish we had read earlier. ##### Two phases, two bottlenecks Prefill takes the whole prompt in one forward pass. Every layer sees a matrix of prompt length by hidden size, so each weight matrix is multiplied by a matrix, and the GPU does what it was built for. For a few thousand prompt tokens the work saturates the tensor cores, and time to first token grows roughly linearly with prompt length because the phase is compute bound. Decode produces one token per step: take the last token, run it through every layer, sample, append, repeat. Each step multiplies every weight matrix by a single vector, or a thin matrix if there is a batch. The arithmetic per step is tiny, but every parameter still has to be read from GPU memory to do it, so the step is dominated not by floating point operations but by moving weights from HBM into the compute units, once per generated token. ##### The roofline argument The roofline model makes this precise. For any kernel, define arithmetic intensity as floating point operations performed per byte moved from memory. A device has a peak compute rate and a peak memory bandwidth, and their ratio is the ridge point. Kernels with intensity below the ridge are bandwidth bound and cannot go faster no matter how much compute is available; kernels above it are compute bound. Decode at batch size one performs about two floating point operations per parameter per token while reading two bytes per parameter in bf16, so its intensity is about one operation per byte. An H100 advertises on the order of a thousand teraflops of dense bf16 compute against roughly 3.35 terabytes per second of memory bandwidth, so its ridge point sits near three hundred operations per byte, two orders of magnitude above where decode lives. The tensor cores idle for more than 99 percent of the step, and the lower bound on time per token is simply the size of the weights divided by bandwidth: a 7 billion parameter model in bf16 is 14 gigabytes, so about 4 milliseconds per token on that card, and no kernel optimisation can beat that floor without reducing the bytes. Batching changes the intensity, not the bytes. If sixteen sequences decode together, the weights are read once and used sixteen times, so intensity rises to about sixteen operations per byte and the step takes roughly the same time as a batch of one. That is why throughput per GPU scales almost linearly with batch size until the ridge is reached, and why the design of a serving system is mostly about keeping the batch full. ```figure llm-inference-internals/prefill-decode Prefill writes the cache for every prompt position in one compute bound pass; decode then reads all of it and appends a single column per step, which makes it memory bound. ``` ##### The KV cache and its arithmetic Attention at each layer needs the key and value vectors of every previous token. Without a cache, generating token n would recompute them for all n previous positions at every layer, and total work would grow quadratically with output length. The KV cache stores them once, so each decode step computes the key and value for the new token only and reads the rest. The arithmetic is worth memorising. Per token, the cache holds one key and one value vector per layer per key-value head, so bytes per token equals 2 times layers times kv_heads times head_dim times the bytes per element. ```python def kv_cache_bytes(layers, kv_heads, head_dim, seq_len, batch=1, dtype_bytes=2): per_token = 2 * layers * kv_heads * head_dim * dtype_bytes return per_token * seq_len * batch def show(name, seq_len, **cfg): per_tok = kv_cache_bytes(seq_len=1, **cfg) total = kv_cache_bytes(seq_len=seq_len, **cfg) print(f"{name}: {per_tok / 2**10:.0f} KiB per token, " f"{total / 2**30:.2f} GiB at {seq_len} tokens") show("7B, 32 layers, MHA 32 heads", 4096, layers=32, kv_heads=32, head_dim=128) show("70B, 80 layers, GQA 8 kv heads", 32768, layers=80, kv_heads=8, head_dim=128) ``` A classic 7 billion parameter model with 32 layers of multi-head attention, 32 heads of dimension 128, holds half a mebibyte of cache per token in bf16. A single 4096 token sequence costs 2 GiB, a seventh of the weights, and sixteen such sequences cost more than the weights themselves. Grouped query attention exists largely because of this arithmetic: sharing each key-value head across several query heads divides the per token cost by the group size, so a 70 billion parameter model with 80 layers and 8 key-value heads holds 320 KiB per token, and a 32k context still costs 10 GiB per sequence. The cache is also read on every step, and at long context that read competes with the weight read. For the 7B model above, a 28k token context means 14 gigabytes of cache, the same as the weights, and the per token latency floor doubles. ##### Paged attention and why fragmentation matters The naive way to hold the cache is to reserve, for each request, a contiguous region sized for the maximum sequence length. Most requests stop well before the maximum, and the reserved space cannot be used by anyone else while the request is alive; that is internal fragmentation. Requests of different lengths also leave gaps between reservations too small to fit a new one, which is external fragmentation. The measurements published with vLLM showed that under this scheme most of the cache memory held nothing useful. Paged attention borrows the answer from operating systems. The cache is split into fixed size blocks of, say, sixteen tokens; a sequence owns a list of blocks that need not be contiguous, and a block table maps its logical positions to physical blocks, exactly like a page table. A sequence takes one block at a time from a free list and releases them the moment it finishes. Waste falls to less than one block per sequence, and blocks holding a shared prefix such as a system prompt can be mapped into many sequences at once, with copy on write when one diverges. This matters because batch size is limited by how many sequences' caches fit in memory after the weights, and every byte recovered from fragmentation is one more use of the weight reads that were going to happen anyway. ##### Static batching versus continuous batching Static batching groups requests up front, pads them to the same length, runs prefill for all of them, and decodes step after step until every sequence has emitted its end token. Sequences finish at different times, so the short ones sit in the batch with their slots occupied but idle until the longest one ends, and requests that arrive meanwhile wait for the next batch. Both waste the resource that decides throughput, and the waiting lands in the time to first token users see. Continuous batching, introduced as iteration-level scheduling in the Orca paper and now standard in vLLM, TGI and TensorRT-LLM, moves the scheduling decision to every decode step. After each step the scheduler releases the cache blocks of finished sequences and admits waiting requests into the freed capacity, and a newcomer's prefill can be folded into the same step as everyone else's decode. ```figure llm-inference-internals/continuous-batching With static batching three of four slots sit idle waiting for the longest sequence; with continuous batching a freed slot is refilled at the very next step and the batch stays full. ``` There is a cost that only shows up in latency histograms. Folding a long prefill into a decode step makes that step as slow as the prefill, so every other sequence sees a spike in inter token latency whenever a large prompt arrives. Chunked prefill caps the prompt tokens processed per step, trading a slightly longer time to first token for the newcomer against a stable cadence for everyone else; the chunk size is one of the few serving parameters we tune per workload. ##### Speculative decoding Batching raises throughput. Speculative decoding attacks latency for a single sequence, and it follows from the roofline: if the target model is going to read all its weights to produce one token, verifying several proposed tokens in one forward pass costs almost the same, because the extra positions add arithmetic to a step with idle compute to spare. A small draft model generates a handful of tokens autoregressively; five is typical. The target then runs once over the prompt plus the draft, which yields its own next-token distribution at every one of those positions in parallel. Each draft token is checked in order, and the first rejection stops the process: everything before it is kept, one corrected token is sampled at the rejection point, and the rest is discarded. If every draft token is accepted, the target's distribution at the final position provides one bonus token for free. The acceptance rule is what makes this exact rather than a heuristic. With draft distribution q and target distribution p at a position, accept the draft token x with probability min(1, p(x) / q(x)). On rejection, sample the replacement from the residual distribution, which is max(0, p minus q) normalised to sum to one. The combination of accepted draft tokens and residual samples is distributed exactly as if the target had sampled alone, so speculative decoding never changes what the model would have said, only how quickly it says it. ```python import numpy as np def speculative_step(p_target, q_draft, draft_tokens, rng): """p_target has one row per draft position plus one extra row for the bonus token; q_draft has one row per draft position. Returns the token ids kept this step.""" kept = [] for i, x in enumerate(draft_tokens): p, q = p_target[i], q_draft[i] ratio = p[x] / max(q[x], 1e-12) if rng.random() < min(1.0, ratio): kept.append(x) continue residual = np.maximum(p-q, 0.0) residual /= residual.sum() kept.append(int(rng.choice(len(p), p=residual))) return kept bonus = p_target[len(draft_tokens)] kept.append(int(rng.choice(len(bonus), p=bonus))) return kept rng = np.random.default_rng(0) vocab, gamma = 50, 5 p = rng.dirichlet(np.ones(vocab) * 0.3, size=gamma + 1) q = 0.7 * p[:gamma] + 0.3 * rng.dirichlet(np.ones(vocab), size=gamma) draft = [int(rng.choice(vocab, p=q[i])) for i in range(gamma)] print(draft, "->", speculative_step(p, q, draft, rng)) ``` The speedup depends on how often the draft agrees with the target, which is high on boilerplate and code and low on creative or numeric content, so the same configuration can double the token rate on one workload and do nothing on another. It also depends on batch size: at a large batch the target step is already compute bound, the verification positions are no longer free, and speculation can make things slower. The draft need not be a separate model either: prompt lookup uses n-grams from the prompt itself, which suits answers that quote retrieved sources, and Medusa and EAGLE attach small prediction heads to the target's own hidden states. ```figure llm-inference-internals/speculative-decoding A draft model proposes five tokens in cheap sequential steps, the target scores all of them in one pass, and the acceptance rule keeps three, resamples the fourth and discards the fifth. ``` ##### Quantisation of weights and of the cache Weight quantisation to eight or four bits, with GPTQ, AWQ or plain round-to-nearest with per-group scales, reduces the bytes read per token, which is the number decode latency is bound by. A 4-bit 7B model is about 4 gigabytes instead of 14, and the per token floor drops by the same factor; the dequantisation arithmetic is paid with compute that was idle anyway. The risk is accuracy, because activation outliers make some channels far more sensitive than others, so we never ship a quantised model without running the same evaluation harness against the unquantised one. KV cache quantisation to fp8 or int8 shrinks the per token cache, which raises the batch size that fits and cuts the cache read at long context. Keys carry a few outlier channels, so per channel scaling for keys and per token scaling for values, the arrangement used by KIVI, holds up better than one scheme for both. This is also where long context accuracy quietly degrades, because an error in an early key affects every later attention score that reads it. ##### What a practitioner should measure Three numbers describe a serving stack. Time to first token is the prefill cost plus any time spent in the queue; it grows with prompt length and is dominated by compute. Inter token latency is the decode cadence; it is dominated by bytes per step and rises with batch size and context length. Tokens per second per GPU is throughput, and it rises with batch size until compute saturates. These trade against each other along the batch size axis, so a single number is never a benchmark. We record the curve, throughput against p50 and p95 inter token latency, for prompt and output lengths drawn from real traffic, and we log queueing time separately from prefill time because the two need different fixes. Goodput, the throughput of requests that met their latency target, is the number a capacity decision should be made on. Decode is bound by bytes, not flops, so anything that reduces bytes per token lowers latency, anything that spreads the weight read across more sequences raises throughput, and speculative decoding buys latency at small batch by spending compute that would otherwise sit idle. Those are the levers, and measured properly each of them shows up in the curve before it shows up in the bill. ### Tiled Inference on Large Rasters Without Seams - Tags: Segmentation, Raster, PyTorch, GIS - Summary: How we ran a field boundary model over district-scale rasters: halos and overlap, Gaussian blending instead of averaging, test time augmentation, a rolling buffer for windowed writes, and a detector that flags seams before anyone opens the output. #### Tiled Inference on Large Rasters Without Seams The field boundary model we built for Kishtwar was trained on fixed-size crops and then had to be run over rasters covering an entire district. That gap between what the network sees in training and what it has to produce in production is where most of the ugly artefacts in raster segmentation come from: faint grid lines in the probability map, parcels that close on one side of a tile and stay open on the other, and boundaries that jitter at exact multiples of the tile size. This note covers how we tiled the district, blended the overlaps, wrote the result back without holding it in memory, and caught the seams we could not see by eye. ##### Why the model cannot see the whole district The arithmetic kills the naive approach before any engineering does. A raster with tens of thousands of pixels on each side, in three or four bands, is already gigabytes as float32 input. The network then multiplies that. HRNet-W48 keeps its highest resolution stream at a quarter of the input resolution with 48 channels, and its head concatenates all four streams at that scale into 720 channels, so the head activations alone are an order of magnitude larger than a four band input tensor, before the intermediate stages, before the batch dimension, before the workspace that cuDNN reserves for convolutions. No card we had access to could hold a district in one forward pass, and even if one could, it would be the wrong thing to do. The second reason is statistical. Batch normalisation statistics, the augmentation policy and the effective receptive field were all shaped by crops of a fixed size, and border padding and fixed-stride downsampling mean a network does not behave identically at a 1024 pixel crop and a 30000 pixel scene. Inference at the training window size keeps the model inside the distribution it was optimised for. Tiling is not a workaround; it is the honest way to use the model. ##### Receptive field and the edge band Every output pixel is computed from a neighbourhood of input pixels, the receptive field. For a deep backbone that neighbourhood is hundreds of pixels wide in theory, but the effective receptive field, the region where changing the input actually changes the output, is much smaller and roughly Gaussian in shape. A pixel in the middle of a tile gets its whole neighbourhood. A pixel on the tile edge is missing half of it, and the missing half has been replaced by whatever the padding provides, usually zeros. The model never saw a field that ends in a wall of zeros during training, so its predictions in a band along each edge are worse than in the middle, and worse in a way that depends on where the tile happened to fall rather than on the terrain. With terraced parcels this shows up brutally. A terrace wall crossing a tile edge is predicted as a boundary in the interior of the tile and fades out as it approaches the edge, because the network can no longer see the ridge continue. Stitch those tiles directly and every boundary stops a few pixels short of every tile border, which vectorises into parcels with gaps exactly where the tiles met. ```figure tiled-inference-large-rasters/halo-window A window is read with a halo on every side so that pixels in the core still see a full receptive field; consecutive windows overlap by twice the halo. ``` ##### Overlap and the halo The fix is to read more than you keep. Each window is read with a halo of extra pixels on every side, wide enough that every pixel we intend to keep has its full effective receptive field inside the window. We set the halo empirically rather than from the layer stack, because the theoretical receptive field of HRNet is far larger than the region that matters. We tried a few widths, looked at how far into the tile the prediction along a straight terrace stayed consistent, and picked the smallest width beyond which the edge band vanished. Overlap changes the window generator in two ways. The stride becomes the tile size minus the overlap. And the last row and column of windows should be snapped to the raster edge rather than padded, because a window that hangs off the raster gets zero padding on that side, which is exactly the edge effect we were trying to remove. Snapping means the final windows overlap their neighbours a little more than the others, which costs almost nothing. ```python from typing import Iterator from rasterio.windows import Window def grid_axis(length: int, tile: int, stride: int) -> list[int]: """Window start offsets along one axis. The last window is snapped to the raster edge.""" if length <= tile: return [0] starts = list(range(0, length-tile+1, stride)) if starts[-1]+tile < length: starts.append(length-tile) return starts def iter_windows(height: int, width: int, tile: int = 1024, overlap: int = 256) -> Iterator[tuple[int, int, Window]]: """Yield (row_index, col_index, window). Neighbouring windows share `overlap` pixels.""" stride = tile-overlap assert 0 < stride < tile, "overlap must be positive and smaller than the tile" rows = grid_axis(height, tile, stride) cols = grid_axis(width, tile, stride) for i, r in enumerate(rows): for j, c in enumerate(cols): yield i, j, Window(col_off=c, row_off=r, width=tile, height=tile) ``` Rasters smaller than a tile in one dimension still need padding. We read those with `boundless=True` and a fill value and clip the padded region when writing back. It is the only case where padding is unavoidable. ##### Blending: why a plain average leaves seams Once windows overlap there are several predictions for every pixel in the overlap zone and a choice about how to combine them. Crop centre is the simplest option: keep the core of each window, discard the halo, never blend. Every kept pixel came from a window where it had full context, but there is still a discontinuity. The pixel on one side of the core border came from window A, its neighbour came from window B, and the two windows disagree slightly everywhere, so a faint step remains along the core grid. Averaging the overlap is the obvious next step and it is worse than it looks. Inside the overlap zone every pixel is the mean of two predictions; one pixel outside the zone is a single prediction. The weight given to window A jumps from one to a half at the overlap boundary. Any systematic difference between the two windows, and there is always one because their contexts differ, becomes a visible line at exactly that boundary. In the Kishtwar probability maps this appeared as pairs of faint parallel lines around each seam, one at each edge of the overlap. Linear feathering ramps the weight from one to zero across the overlap. The weight is now continuous but its derivative is not, and a sharp model still leaves a softer trace at both ends of the ramp. A Gaussian weight map removes the remaining trace because it is smooth everywhere, including at the window edge where it approaches, but never reaches, zero. It also does something the others do not: it trusts the centre of a window more than its edge, which is the right prior given how the edge band behaves. The map is the outer product of a one-dimensional Gaussian with itself, normalised to peak at one, with a small floor so the sum of weights is never zero at a pixel covered by only one window. ```figure tiled-inference-large-rasters/weight-blend Plain averaging gives window A a weight that drops from one to a half at the overlap edge; a Gaussian weight decays smoothly, so the two predictions cross-fade without a step. ``` ```python import numpy as np def gaussian_weight(tile: int, sigma_scale: float = 0.125, floor: float = 1e-3) -> np.ndarray: """2D weight peaking at the tile centre. sigma of tile/8 is the value nnU-Net uses.""" axis = np.arange(tile, dtype=np.float32) centre = (tile-1)/2.0 sigma = tile*sigma_scale g = np.exp(-((axis-centre)**2)/(2.0*sigma**2)) w = np.outer(g, g) w /= w.max() return np.maximum(w, floor).astype(np.float32) ``` We blend probabilities after the softmax rather than raw logits. Logits from two windows can sit on different scales, and averaging them before the softmax does not give the average of the two opinions. Probabilities are bounded, additive, and produce the argmax you expect. ##### Test time augmentation across tiles Flipping and transposing each window, running the model on every variant, undoing the transform on the output and averaging is the cheapest accuracy we ever bought. It matters more for boundaries than for regions, because a boundary detector is asymmetric in ways that show up as a half pixel bias in one direction, and averaging over the flips cancels that bias. A square tile has eight dihedral variants; we used four (identity, horizontal flip, vertical flip, transpose), because the other four doubled the compute for a change we could not see. TTA interacts with blending in one useful way. Averaging over augmentations reduces the variance of each window's prediction, so neighbouring windows disagree less, so seams are fainter before any weighting is applied. It does nothing about the discontinuity in the weight map, which is why it complements the Gaussian map rather than replacing it. ```python import torch @torch.no_grad() def predict_batch(model: torch.nn.Module, x: torch.Tensor, tta: bool = True) -> np.ndarray: """x: (N, C, T, T) normalised bands on the GPU. Returns class probabilities (N, K, T, T) on the CPU.""" with torch.autocast("cuda", dtype=torch.float16): p = model(x).softmax(1) if tta: p = p+model(x.flip(-1)).softmax(1).flip(-1) p = p+model(x.flip(-2)).softmax(1).flip(-2) p = p+model(x.transpose(-1, -2)).softmax(1).transpose(-1, -2) p = p/4.0 return p.float().cpu().numpy() ``` ##### Memory planning and batching windows Three memory budgets have to be planned separately: the GPU, the accumulators, and the output file. On the GPU the constraint is the batch of windows times the TTA multiplier. Half precision autocast halves activation memory and, on the cards we used, roughly doubled throughput with no change we could measure in the boundaries. We sized the batch by trial on the largest window and then left it alone. Dynamic batching by free memory sounds attractive and in practice produces out of memory errors at three in the morning when a different process wakes up on the same machine. Reads matter too: a GeoTIFF with internal tiling and a GDAL block cache sized for a row of windows keeps the GPU from waiting on disk. The accumulators are the real problem on large rasters. A weighted sum and a weight sum for every class over the whole raster in float32 is several times the size of the input and does not fit on a modest machine. The observation that saves us is that a window in grid row i never touches raster rows above the start of grid row i, so once every window in row i has been added, all raster rows above the start of row i plus one are final. We keep a rolling buffer two tiles tall, add windows into it, and after each grid row flush the finished rows to disk and shift the buffer up. Peak memory then depends on the raster width and the tile size, not on the raster height. The output is a tiled GeoTIFF with internal blocks and deflate compression, with BIGTIFF allowed when the driver decides it is needed, so that the vectorisation step afterwards can read it in windows as well. ##### Writing back with rasterio windowed writes Rasterio will write any rectangular window of a dataset opened in write mode as long as the profile was fixed at creation. The rolling buffer flushes one band of complete rows per grid row with a single write call. ```python from itertools import groupby import numpy as np import rasterio from rasterio.windows import Window class RollingWriter: """Accumulates weighted probabilities in a buffer two tiles tall and flushes rows no later window can touch.""" def __init__(self, dst: rasterio.io.DatasetWriter, n_classes: int, tile: int): self.dst, self.tile, self.top = dst, tile, 0 self.acc = np.zeros((n_classes, tile*2, dst.width), np.float32) self.wsum = np.zeros((1, tile*2, dst.width), np.float32) def add(self, win: Window, prob: np.ndarray, weight: np.ndarray) -> None: r0, c0 = int(win.row_off)-self.top, int(win.col_off) h = min(int(win.height), self.dst.height-int(win.row_off)) w = min(int(win.width), self.dst.width-c0) self.acc[:, r0:r0+h, c0:c0+w] += (prob*weight)[:, :h, :w] self.wsum[:, r0:r0+h, c0:c0+w] += weight[:, :h, :w] def flush_until(self, row: int) -> None: """Write raster rows [self.top, row), then roll the buffer up so row becomes its first line.""" n = row-self.top if n <= 0: return out = self.acc[:, :n]/np.maximum(self.wsum[:, :n], 1e-6) self.dst.write(out, window=Window(0, self.top, self.dst.width, n)) self.acc = np.roll(self.acc, -n, axis=1) self.wsum = np.roll(self.wsum, -n, axis=1) self.acc[:, -n:] = 0.0 self.wsum[:, -n:] = 0.0 self.top = row def segment_raster(src_path: str, dst_path: str, model: torch.nn.Module, mean: torch.Tensor, std: torch.Tensor, tile: int = 1024, overlap: int = 256, batch: int = 4, n_classes: int = 2) -> None: weight = gaussian_weight(tile)[None] with rasterio.open(src_path) as src: profile = src.profile.copy() profile.update(count=n_classes, dtype="float32", nodata=None, tiled=True, blockxsize=512, blockysize=512, compress="deflate", BIGTIFF="IF_SAFER") rows = grid_axis(src.height, tile, tile-overlap) with rasterio.open(dst_path, "w", **profile) as dst: writer = RollingWriter(dst, n_classes, tile) pending: list[tuple[Window, np.ndarray]] = [] def run_pending() -> None: x = torch.from_numpy(np.stack([t for _, t in pending])).pin_memory().cuda(non_blocking=True) probs = predict_batch(model, (x-mean)/std) for (win, _), p in zip(pending, probs): writer.add(win, p, weight) pending.clear() for i, group in groupby(iter_windows(src.height, src.width, tile, overlap), key=lambda t: t[0]): for _, _, win in group: pending.append((win, src.read(window=win, boundless=True, fill_value=0).astype(np.float32))) if len(pending) == batch: run_pending() if pending: run_pending() writer.flush_until(rows[i+1] if i+1 < len(rows) else src.height) ``` The `mean` and `std` tensors are the per-band statistics saved with the checkpoint, shaped `(1, C, 1, 1)` and already on the GPU. Two details caused real bugs before they were right: rasterio window offsets are floats and must be cast to `int` before slicing, and the flush after each grid row has to use the start of the next grid row rather than `top` plus the stride, because the snapped last row breaks the regular stride. ```mermaid flowchart LR A[Window with halo] --> B[Batch of windows] B --> C[Model plus TTA] C --> D[Gaussian weight] D --> E[Rolling buffer] E --> F[Windowed write] F --> G[Seam ratio check] ``` ##### Detecting seam artefacts automatically Looking for seams by eye on a district does not scale, and the faint ones only show when the probability map is stretched, so we wrote a detector and ran it after every inference job. It uses the one thing we know for certain: where the window edges are. Take the gradient magnitude of the boundary class probability with a Sobel filter, average it in a narrow band along every window edge, where the weighting changes, and divide by the same statistic on control lines half a stride away, where no edge lies. Terrain contributes equally to both, so a ratio near one means the edges are indistinguishable from the rest of the map, and a ratio well above one flags seams. One number per job goes straight into the job log and the alert threshold. ```figure tiled-inference-large-rasters/seam-detector The detector compares the gradient of the probability map in a narrow band along window edges against the same statistic on control lines half a stride away; a ratio near one means no seam. ``` ```python from scipy import ndimage def seam_ratio(prob: np.ndarray, starts: list[int], tile: int, band: int = 2) -> float: """Ratio of gradient magnitude on vertical window edges to gradient on control columns. prob: (H, W) boundary probability for one stripe; starts: column offsets from grid_axis. Call again with prob.T and the row offsets for horizontal seams.""" grad = np.hypot(ndimage.sobel(prob, axis=1), ndimage.sobel(prob, axis=0)) width = prob.shape[1] edges = sorted({s for s in starts if s > 0} | {s+tile for s in starts if s+tile < width}) stride = starts[1]-starts[0] if len(starts) > 1 else tile controls = [e+stride//2 for e in edges if e+stride//2+band < width] def mean_on(cols: list[int]) -> float: return float(np.mean([grad[:, c-band:c+band+1].mean() for c in cols])) return mean_on(edges)/max(mean_on(controls), 1e-6) ``` The detector runs on full resolution stripes of the output, never on a downsampled overview, because a seam is a one or two pixel feature and any resampling averages it away. We also kept a second check after vectorisation: a histogram of polygon vertex coordinates modulo the stride. Seams that survive into the vector layer show up as a spike at zero, because parcel edges snap to the grid where the boundary probability stepped. The ratio earned its place the first time the imagery changed. A delivery at a slightly different ground resolution made the halo marginal, and the ratio climbed on the first job before anyone had opened the output in QGIS. Widening the halo brought it back down, which is why the halo is a configuration value rather than a constant in the code. ### Inside HNSW: How Approximate Nearest Neighbour Search Trades Recall for Latency - Tags: Vector Search, HNSW, pgvector, Algorithms - Summary: How the layered graph behind pgvector and hnswlib routes a query, what M, efConstruction and efSearch each change, why deletes and filters break the usual assumptions, how IVF and product quantisation compare, and how to measure recall against brute force before trusting an index in production. #### Inside HNSW: How Approximate Nearest Neighbour Search Trades Recall for Latency The retrieval half of the RAG pipeline we run for a private knowledge base lives in PostgreSQL with pgvector, and for a long time we treated the index as a black box. Then a filtered query returned nothing for a tenant who certainly had documents, and a recall check run mostly out of curiosity showed the index missing passages that a brute force scan found without effort. Understanding what the index actually does turned out to be the difference between a retrieval layer we could reason about and one we could only restart. ##### Why exact search stops scaling Brute force nearest neighbour search is a dot product or a distance between the query and every vector in the corpus, so its cost is the corpus size times the dimension. Ten million embeddings of 1536 float32 dimensions occupy about 61 gigabytes, and every query has to stream all of it through the processor. The scan is bound by memory bandwidth rather than arithmetic, and no amount of vectorised code changes the number of bytes. The classic answer for low dimensional data, a tree that splits space along coordinates and prunes branches that cannot contain a closer point, fails at embedding dimensions. As dimension grows, distances between random points concentrate around the same value, the pruning bound almost never excludes a branch, and the tree degenerates into a scan with extra pointer chasing. So the field gave up exactness. An approximate index returns neighbours that are among the true closest with high probability, and it exposes a knob that trades that probability against query time. ##### Navigable small worlds The idea behind HNSW is a graph. Every vector is a node, and each node keeps links to a small number of near neighbours. To search, start at some node, look at its neighbours, move to whichever is closest to the query, and repeat until no neighbour is closer than where you stand. This greedy routing needs the graph to have two properties at once: short links so the final approach is precise, and long links so the walk from a distant entry does not take thousands of hops. Graphs with both are called navigable small worlds, and Kleinberg showed that with the right distribution of long links greedy routing takes a number of hops that grows only logarithmically with the number of nodes. The original NSW construction produced long links by accident of ordering: points were inserted in random order and linked to their nearest neighbours among the points already present, so early points reached across the whole space and later points filled in the local structure. It worked, but long and short links lived in the same neighbour lists, so a walk that needed a long hop had to evaluate all the short ones at every node too. ##### Layers and the greedy descent HNSW separates the scales into layers. When a vector is inserted it is assigned a maximum level drawn from a geometric distribution: level equals the floor of negative ln(u) times mL, where u is uniform on (0, 1) and mL is 1 over ln(M). The vector then appears in layer 0 and in every layer up to its level. Layer 0 holds everything, layer 1 holds a fraction of it, layer 2 a fraction of that, so the upper layers are sparse and their links are necessarily long. Search enters at the single entry point on the top layer and runs greedy routing there until no neighbour is closer. It then drops to the layer below at that same node and repeats. Each layer refines the position, like a skip list. At layer 0 the search widens from a single walker to a beam: it keeps the efSearch best candidates seen so far, expands them in order of distance, and stops when the closest unexpanded candidate is farther than the worst result already held. The top k of that beam is the answer. ```figure hnsw-vector-index-internals/layered-descent Greedy routing hops across the sparse top layer, drops to the next layer at its best node, and only widens into a beam of efSearch candidates at layer 0. ``` The whole procedure fits in a page of Python over adjacency dictionaries, one per layer, and is the shape of what hnswlib, faiss and pgvector implement in C. ```python import heapq import numpy as np def search_layer(vectors, graph, query, entry, ef): """graph maps a node id to its neighbour ids on one layer. Returns up to ef (distance, node) pairs, closest first.""" dist = lambda i: float(np.linalg.norm(vectors[i]-query)) d0 = dist(entry) visited = {entry} candidates = [(d0, entry)] # min-heap: closest unexpanded first results = [(-d0, entry)] # max-heap on distance: worst result on top while candidates: d, node = heapq.heappop(candidates) if d > -results[0][0] and len(results) >= ef: break for nb in graph[node]: if nb in visited: continue visited.add(nb) dn = dist(nb) if len(results) < ef or dn < -results[0][0]: heapq.heappush(candidates, (dn, nb)) heapq.heappush(results, (-dn, nb)) if len(results) > ef: heapq.heappop(results) return sorted((-d, n) for d, n in results) def hnsw_search(vectors, layers, query, entry, ef_search, k): """layers[0] is the base layer; layers[-1] is the top layer.""" node = entry for graph in reversed(layers[1:]): node = search_layer(vectors, graph, query, node, ef=1)[0][1] return search_layer(vectors, layers[0], query, node, ef_search)[:k] ``` ##### M, efConstruction and efSearch Three parameters control the index. M is the number of links each node keeps on the upper layers; layer 0 keeps twice as many. pgvector calls it m and defaults to 16. A larger M gives every node more directions to route in, which raises recall on high dimensional data and on clustered data, at the cost of more memory per node and more distance computations per hop. The memory is usually not the vectors' rival: at 1536 dimensions a float32 vector is 6 KiB while 32 links are 128 bytes. The distance computations are the real cost, because every hop evaluates all of a node's neighbours. The way those M links are chosen matters as much as their number. The naive choice, the M nearest candidates, tends to point every link into the same dense cluster. HNSW's heuristic instead accepts a candidate only if it is closer to the node than to any neighbour already accepted, which spreads links across directions and keeps separate clusters connected. Indexes built without it have islands that greedy routing cannot reach. efConstruction is the beam width used during insertion, when the index searches for each new vector's neighbours. A wider beam finds better neighbours and produces a better graph, so recall at a given efSearch goes up, but build time grows almost linearly with it; pgvector defaults to 64. efSearch is the beam width at query time and the only parameter that can be changed without rebuilding. It must be at least k, and in pgvector the setting hnsw.ef_search caps the number of rows an index scan returns, so a LIMIT above it is silently truncated. Raising efSearch widens the candidate frontier, which recovers neighbours that a narrow beam walked past, with diminishing returns and roughly linear latency. ```figure hnsw-vector-index-internals/ef-search A narrow beam expands few nodes and walks past two of the five true neighbours; a wider beam visits more of the graph, finds all five, and pays for it in distance computations. ``` ##### Why deletes are hard Insertion is a search followed by linking; deletion has no such symmetry. Removing a node removes its links, and any node that was reachable only through it becomes unreachable, so a graph that has seen many deletions slowly loses the property the search depends on. The standard answer is a tombstone: the deleted node stays in the graph for routing and is excluded from results. hnswlib does this with mark_deleted, and pgvector does it by leaving the dead tuple in the index until vacuum removes it and repairs the neighbour lists of the nodes that pointed at it. The cost is that tombstones consume the beam. A search with efSearch of 40 that passes through twenty dead nodes has an effective beam of twenty, and recall degrades in proportion to the churn. Updates are a delete and an insert, so re-embedding a corpus with a new model is a full rebuild whether or not it is called one. On a corpus that changes constantly we vacuum the table aggressively and rebuild the index concurrently on a cadence tied to the fraction of tuples that have changed. ##### Filtered search and the pre versus post filter trap Almost every real query carries a predicate: this tenant, this document type, this date range. There are two orders in which to combine a predicate with an approximate search, and choosing the wrong one produces the empty result that started this note. Post filtering runs the approximate search first and applies the predicate to its results. If the predicate matches one percent of the corpus, the expected number of matches among the ten nearest neighbours is a tenth of a row, so the query returns nothing while the tenant has thousands of documents. Pre filtering evaluates the predicate first and searches only the survivors. When the predicate is selective this is cheap and exact, but when it matches half the corpus it degenerates into a brute force scan over millions of rows. In pgvector the planner chooses. With a selective predicate and a B-tree index on the filter column it may pick a bitmap scan and compute exact distances on the survivors, which is pre filtering. With an ordered index scan on the HNSW index it fetches hnsw.ef_search candidates by distance, applies the predicate, and returns what survives, which is post filtering and can return fewer rows than the LIMIT. The tools that turn this into something reliable are raising ef_search so that enough candidates survive, keeping table statistics fresh so the selectivity estimate is right, partial indexes when the predicate has few values such as a tenant column, and iterative scans, which pgvector added so that the index keeps pulling candidates from the graph until the LIMIT is satisfied or a scan budget is exhausted. ```sql SET hnsw.ef_search = 200; SET hnsw.iterative_scan = relaxed_order; SELECT id, embedding <=> $1 AS distance FROM chunks WHERE tenant_id = $2 ORDER BY embedding <=> $1 LIMIT 10; ``` The general solution, filtering inside the graph traversal, is what engines such as Qdrant, Vespa and filtered DiskANN do: the predicate is evaluated during the walk, and extra links are built so that the subgraph satisfying common predicates stays connected. Inside PostgreSQL we do not have that, so we benchmark every predicate shape we serve rather than assuming the unfiltered recall carries over. ```figure hnsw-vector-index-internals/filter-trap Post filtering fetches the ten nearest rows and then applies the predicate, which leaves nothing; pre filtering restricts to the tenant first and then orders by distance. ``` ##### IVF and product quantisation, and when they win HNSW is not the only approximate index, and the alternatives trade differently. Inverted file indexes cluster the corpus with k-means into a few thousand cells. A query is compared against the centroids, the nprobe nearest cells are opened, and only their members are scanned. There is no graph, so memory is just the vectors plus a cell assignment, and build is a single clustering pass. The weakness is that the centroids are fixed at build time: data inserted later is assigned to the nearest existing cell, and as the distribution drifts recall falls. pgvector's ivfflat documentation says to build after the table is loaded for this reason, and in a corpus with steady ingest we found HNSW's incremental insertion worth its extra memory. Product quantisation attacks the bytes rather than the candidate count. Each vector is split into subvectors, each subvector is replaced by the index of its nearest centroid in a small codebook, and a 6 KiB embedding becomes a few dozen bytes whose distances are computed with lookup tables. Recall drops because the codes are lossy, and the usual remedy is to rerank the best few hundred candidates with their full vectors. faiss combines both ideas as IVF-PQ for billion scale corpora, and pgvector offers a cheaper version of the same trade in half precision and binary vectors with an exact rerank. For a corpus that fits in memory and a latency budget in milliseconds, HNSW over full vectors is the right default; the compressed indexes earn their complexity when the corpus does not fit. ##### Benchmark recall before trusting the index None of the parameters above mean anything without a measurement, and the measurement is simple. Recall at k is the fraction of the true k nearest neighbours that the index returned, averaged over a set of queries, and the truth comes from brute force, which is slow but only runs once per benchmark set. Two details decide whether the number is honest. The queries must come from the real query distribution, not from vectors sampled out of the corpus, because a question embedding and a passage embedding live in different regions of the space and recall on corpus vectors flatters the index. The benchmark must also run with the same predicates production uses, because recall under a filter is a different quantity from recall without one. ```python import time import numpy as np def brute_force_topk(corpus, queries, k): """Rows of corpus and queries are unit vectors, so a larger dot is closer.""" sims = queries @ corpus.T return np.argsort(-sims, axis=1)[:, :k] def recall_at_k(approx_ids, exact_ids): k = exact_ids.shape[1] hits = [len(set(a[:k]) & set(e)) for a, e in zip(approx_ids, exact_ids)] return float(np.mean(hits)) / k def benchmark(search_fn, corpus, queries, k, ef_values): exact = brute_force_topk(corpus, queries, k) for ef in ef_values: ids, lat = [], [] for q in queries: t0 = time.perf_counter() ids.append(search_fn(q, k, ef)) lat.append(time.perf_counter()-t0) lat_ms = np.array(lat) * 1e3 print(f"ef={ef:4d} recall@{k}={recall_at_k(np.array(ids), exact):.3f} " f"p50={np.percentile(lat_ms, 50):.2f} ms p95={np.percentile(lat_ms, 95):.2f} ms") ``` The search_fn for pgvector is a query with SET LOCAL hnsw.ef_search inside a transaction, run against the real table with the real predicate. We sweep efSearch, record recall against p50 and p95 latency, and pick the smallest value on the flat part of the recall curve. The sweep is rerun after any large ingest, after a vacuum that removed many tuples, and whenever the embedding model changes, because each of those moves the curve. Once the number was on a dashboard, the empty result for that tenant became a line on a graph we could see coming. ### Boundary Aware Losses for Thin Structure Segmentation - Tags: Deep Learning, Loss Functions, Segmentation - Summary: Why pixel cross-entropy and Dice fail on one to two pixel boundaries, how distance weighting and the signed distance boundary loss fix it, and why the boundary F score with a tolerance band is the metric to trust. #### Boundary Aware Losses for Thin Structure Segmentation The training target for the Kishtwar parcel work was a boundary line one to two pixels wide at the imagery resolution, rasterised from digitised polygons. The first HRNet we trained on it with plain pixel cross-entropy reported excellent pixel accuracy and produced a boundary map that was almost empty: a few strong terrace walls came through, the thin bunds between neighbouring terraces did not. This note is about why the loss was the problem, which replacements earned a place in the final objective, and the metrics that told the truth when accuracy did not. ##### Why pixel cross-entropy under-weights thin boundaries Cross-entropy is an average over pixels, and a thin boundary is a tiny fraction of them. A parcel roughly n pixels across has about n squared interior pixels and about 4n boundary pixels, so the boundary class is a share of about 4/n of the image: for a parcel 64 pixels wide, around six percent at one pixel width, and still a small minority for the much smaller Kishtwar terraces. The gradient the network receives is dominated by background pixels, and a model that predicts background everywhere with a mild hedge reaches a loss that looks respectable. The usual first response, inverse frequency class weights, rebalances the sum but introduces its own failure. Once a boundary pixel is worth twenty background pixels, every pixel next to the true line is cheap to get wrong in the boundary direction, and the predicted lines fatten. When a fat line is thinned back to one pixel for vectorisation it wanders, and two parcels that share a boundary end up with two different edges. The deeper flaw is that cross-entropy has no notion of distance. A prediction one pixel to the left of the true line is exactly as wrong as one in the middle of the field. That is a bad match for a target whose position is itself uncertain by a pixel, because a line rasterised from a polygon outline lands in a different pixel depending on the rasterisation rule. ```figure boundary-aware-losses/pixel-imbalance Along one row of the label raster the boundary is two pixels out of twenty-four; plain cross-entropy gives every pixel one vote, so the summed gradient is almost entirely about the background. ``` ##### Dice and where it fails on tiny structures Soft Dice divides twice the overlap by the sum of the predicted and true set sizes, so it is indifferent to how much background there is. For a one pixel line it misbehaves in three ways we hit directly. The gradient per pixel scales inversely with the set sizes, so a handful of pixels carry an enormous, noisy gradient, and a crop with no boundary in its label produces a degenerate term unless the smoothing constant is large. Computing Dice over the whole batch rather than per crop stabilised this more than any choice of smoothing constant. Dice also has the same blindness to distance, only sharper. A one pixel line predicted one pixel off has zero overlap, the same score as a line on the other side of the field, and the loss offers no gradient in the direction of closer. What it rewards is a fatter prediction, because a thick line overlapping the true one gains more intersection than it loses to the size penalty. Cross-entropy plus Dice gave boundaries that were connected but fat and wobbly, an improvement on empty, not a result we could vectorise. ##### Distance transform weighting The first fix that moved the needle is borrowed from the original U-Net paper, where a per pixel weight map forces the network to learn the thin gaps between touching cells. Our variant is simpler. Take the Euclidean distance transform of the background, so every pixel knows how far it is from the nearest boundary pixel, and build a weight that is a Gaussian bump around the line on top of a constant floor. Pixels on and next to the line dominate the loss, and pixels in the middle of a field keep a small weight so the model still learns to stay quiet there. The width of the bump should follow the label quality rather than the line width. A narrow bump punishes a prediction one pixel off very hard, which is right when the labels are accurate to a pixel and wrong when they are not; when we moved to a reference set digitised at a coarser zoom, the same sigma made the loss jump around and we widened it. The map is cheap, so we compute it in the dataloader workers after augmentation rather than storing it, for a reason covered below. ```python import numpy as np import torch import torch.nn.functional as F from scipy import ndimage def distance_weights(line: np.ndarray, sigma: float = 3.0, w0: float = 8.0, floor: float = 1.0) -> np.ndarray: """Per-pixel weight: `floor` everywhere plus a Gaussian bump around boundary pixels (line == 1).""" dist = ndimage.distance_transform_edt(line == 0) return (floor+w0*np.exp(-(dist**2)/(2.0*sigma**2))).astype(np.float32) class DistanceWeightedCE(torch.nn.Module): def __init__(self, class_weight: list[float] | None = None, ignore_index: int = 255): super().__init__() self.register_buffer("class_weight", None if class_weight is None else torch.tensor(class_weight)) self.ignore_index = ignore_index def forward(self, logits: torch.Tensor, target: torch.Tensor, pixel_weight: torch.Tensor) -> torch.Tensor: """logits (N, C, H, W); target (N, H, W) long; pixel_weight (N, H, W) from distance_weights.""" ce = F.cross_entropy(logits, target, weight=self.class_weight, ignore_index=self.ignore_index, reduction="none") w = pixel_weight*(target != self.ignore_index).float() return (ce*w).sum()/w.sum().clamp_min(1.0) ``` ##### Boundary loss and the signed distance map Distance weighting makes the loss care more about pixels near the line, but the penalty for a wrong pixel still does not grow with how far away it is. The boundary loss of Kervadec and colleagues does exactly that. Build a signed distance map of the ground truth, negative inside the object, positive outside, zero on the boundary, and take the mean over pixels of that map multiplied by the softmax probability of the foreground class. A false positive far from the object costs more than one next to it, and a confident prediction inside the object is rewarded, since the product is negative there. The paper derives this as a first order approximation of the area between the predicted and true boundaries. ```figure boundary-aware-losses/signed-distance The signed distance map is negative inside the parcel and positive outside; multiplying it by the softmax output penalises foreground probability in proportion to its distance from the true edge. ``` The subtlety we ran into is that the loss assumes the foreground has an inside. Our foreground was a line. The interior of a one pixel line is the line itself, so the signed distance is zero on the line and positive everywhere else, and the loss is minimised by predicting nothing at all. Applied to the line channel it pushed the model straight back to the empty map we started with. The fix changed what the network predicts. We gave it two output channels, the parcel interior as a filled region and the boundary as a line, and put each loss where its assumptions hold. The region channel has a proper inside, so it takes the boundary loss together with Dice and cross-entropy. The line channel takes the distance weighted cross-entropy. At inference the morphological gradient of the region prediction gives a closed, one pixel line that agrees with the interior by construction, and the line channel breaks the ambiguous cases where two parcels touch and the region prediction merges them. The two heads share every feature but the last convolution. ```python def signed_distance_map(region: np.ndarray, clip: float = 64.0) -> np.ndarray: """Negative inside the region, positive outside, zero on the boundary; clipped and scaled to [-1, 1].""" fg = region.astype(bool) if not fg.any() or fg.all(): return np.zeros(region.shape, np.float32) outside = ndimage.distance_transform_edt(~fg) inside = ndimage.distance_transform_edt(fg) sdm = outside*(~fg)-(inside-1.0)*fg return (np.clip(sdm, -clip, clip)/clip).astype(np.float32) class BoundaryLoss(torch.nn.Module): """Kervadec et al. 2019: mean over pixels of phi(p) * s(p) for the region class.""" def __init__(self, fg_index: int = 1): super().__init__() self.fg_index = fg_index def forward(self, logits: torch.Tensor, sdm: torch.Tensor) -> torch.Tensor: """logits (N, C, H, W); sdm (N, H, W) from signed_distance_map. Can be negative; that is expected.""" s = logits.softmax(1)[:, self.fg_index] return (s*sdm).mean() ``` Two details matter. The `inside-1` term follows the reference implementation, so object pixels touching the boundary get a distance of zero, which keeps the map antisymmetric across the edge. The clip is ours: a raw distance can reach hundreds of pixels in the middle of a large field, and without it one spurious blob far from any parcel dominates the gradient of the whole batch. Scaling to the unit interval also makes the loss magnitude comparable to Dice, which the schedule below relies on. ##### Hausdorff surrogates The Hausdorff distance is the worst case gap between two boundaries, the number a GIS analyst is implicitly complaining about when a parcel edge is thirty metres out. Karimi and Salcudean proposed training against it with a loss built from distance transforms of both prediction and label, and we tried it. The prediction's distance transform has to be recomputed on the CPU every step, and because the loss is dominated by the few worst pixels, a single mislabelled parcel would swing a whole batch. We kept it as a validation metric instead, the 95th percentile Hausdorff distance between vectorised prediction and reference, where it is informative and free. ##### Combining losses with a schedule The final objective for the region channel is a convex combination of cross-entropy plus Dice on one side and the boundary loss on the other, with the boundary weight ramped from zero to a cap over the first part of training and held there. Kervadec's paper keeps increasing the weight until the boundary term dominates; we stopped short of that. The ramp exists because of how the loss behaves at initialisation. With a near uniform softmax, the gradient is driven by the huge number of far away background pixels, all pushing foreground probability down, and it drives the network to an empty prediction before the region terms have given it a shape to refine. Ramping in after the rough shape exists avoids the collapse; in practice the boundary loss sharpens an existing edge rather than finding one. The cap exists because once the boundary term dominates, the cross-entropy gradient that keeps the interior filled becomes noise. ```python class RegionLineObjective: """Region head: (1-alpha) * (CE + Dice) + alpha * boundary. Line head: distance weighted CE. alpha ramps linearly to alpha_max over ramp_epochs, then holds.""" def __init__(self, ce, dice, boundary, line_ce, alpha_max: float = 0.4, ramp_epochs: int = 30): self.ce, self.dice, self.boundary, self.line_ce = ce, dice, boundary, line_ce self.alpha_max, self.ramp_epochs = alpha_max, ramp_epochs def alpha(self, epoch: int) -> float: return self.alpha_max*min(1.0, epoch/self.ramp_epochs) def __call__(self, region_logits, line_logits, region_target, sdm, line_target, line_weight, epoch: int): a = self.alpha(epoch) region = (1.0-a)*(self.ce(region_logits, region_target)+self.dice(region_logits, region_target)) region = region+a*self.boundary(region_logits, sdm) line = self.line_ce(line_logits, line_target, line_weight) return region+line ``` ##### Metrics that actually measure boundaries None of the above would have been chosen correctly if we had kept validating on mean intersection over union. On a one to two pixel line, IoU measures rasterisation luck. A prediction offset by one pixel scores zero and is a perfect result for the vector product. A fat prediction three pixels wide scores a respectable IoU and is useless, because thinning it produces a wandering line and vectorising it produces slivers between parcels. Pixel accuracy is worse; it was the metric that congratulated us on an empty map. The metric that matched what we saw in QGIS is the boundary F score with a tolerance band, as used in video segmentation benchmarks. Thin both prediction and reference to one pixel, count precision as the fraction of predicted boundary pixels within tau pixels of the reference and recall as the fraction of reference pixels within tau of the prediction, and take the harmonic mean. Tau should come from the annotation uncertainty, not from what makes the number look good; two pixels matched how the polygons had been digitised. We reported the score at two tolerances: improvement at the tighter one means the model is getting more precise, improvement only at the looser one means it is finding edges it used to miss, and the two need different fixes. ```figure boundary-aware-losses/tolerance-band Precision counts predicted boundary pixels that fall inside the tolerance band around the reference, recall counts reference pixels within reach of the prediction; two thin lines one pixel apart score zero IoU and a perfect boundary F. ``` ```python from skimage.morphology import skeletonize def boundary_f_score(pred_line: np.ndarray, gt_line: np.ndarray, tol: float = 2.0) -> tuple[float, float, float]: """Both inputs are binary boundary rasters; they are thinned to one pixel before comparison. Returns (precision, recall, f) with distances measured in pixels.""" pred = skeletonize(pred_line.astype(bool)) gt = skeletonize(gt_line.astype(bool)) dist_to_gt = ndimage.distance_transform_edt(~gt) dist_to_pred = ndimage.distance_transform_edt(~pred) precision = float((dist_to_gt[pred] <= tol).mean()) if pred.any() else 0.0 recall = float((dist_to_pred[gt] <= tol).mean()) if gt.any() else 0.0 f = 0.0 if precision+recall == 0.0 else 2.0*precision*recall/(precision+recall) return precision, recall, f ``` After vectorisation we add object level counts, because the F score cannot see a missing bund that merges two parcels into one. Matching predicted polygons to reference polygons at an overlap threshold and counting the unmatched on each side gives over-segmentation and under-segmentation rates, and those two numbers were what the client's review team reacted to. ##### Practical training notes Label thinning and fixed width. Rasterising polygon outlines gives lines whose width varies with the angle of the edge and the rasterisation rule. We skeletonised every line label to one pixel and then dilated it to a fixed width, the same width throughout training and evaluation. Mixing widths across the dataset is label noise the distance weighted loss punishes hard. Morphological consistency. The line label must be derived from the region label with the same morphological gradient the inference code applies to the region prediction, not rasterised separately from the polygons. When they were produced separately the two disagreed by a pixel along whole sides of parcels, the two heads were trained to contradict each other, and the merged output had doubled edges. Augmentation. Labels are resampled with nearest neighbour only. Arbitrary rotations break a one pixel line into dashes, so geometric augmentation is limited to flips and quarter turns. Distance and weight maps are computed after augmentation from the augmented label, never warped with it, because a warped distance map is no longer a distance map. Ignore regions. Where the reference polygons were known to be unreliable, under tree canopy and at the edges of the digitised area, we rasterised an ignore band and excluded it from every loss. Without it the model was taught confidently wrong boundaries exactly where it had the least visual evidence. Monitoring. We logged the mean width of the thresholded line prediction every validation epoch. It caught the fattening from over-weighted class terms long before the boundary F score made the cause obvious. If a boundary model looks bad, check the labels, then the metric, then the loss, in that order. In our case all three were wrong, and fixing them in the reverse order would have wasted most of the time we had. ### Where the GPU Memory Goes: Fitting a Large Model on One Card - Tags: PyTorch, GPU, Training, Systems - Summary: The memory budget of a training step: weights, gradients, the two Adam moments, activations and workspace. How bf16, loss scaling, checkpointing, accumulation, channels_last and fused optimisers move each term, and how to find the real peak with the PyTorch memory snapshot. #### Where the GPU Memory Goes: Fitting a Large Model on One Card The first out-of-memory error on a training run usually arrives at the worst moment: the model is defined, the loader works, the first batch goes in, and the backward pass dies trying to allocate a few hundred megabytes on a card that already looked full. The reflex is to halve the batch size and try again. That works until it does not, and it teaches nothing about where the memory went. This note is the accounting we do before a run, how each standard technique moves the numbers, and how to read the PyTorch memory snapshot when the arithmetic and the card disagree. ##### The budget before a single activation Training memory splits into two kinds. The first is proportional to the parameter count and lives for the whole run: weights, gradients and optimiser state. The second is proportional to batch size, resolution and depth, is born in the forward pass and dies during the backward: activations. For a model with P trainable parameters in plain fp32 with Adam or AdamW, the persistent part is 4P bytes of weights, 4P bytes of gradients once the first backward has run, and 8P bytes of optimiser state, because Adam keeps two moments per parameter, a running mean of the gradient and a running mean of its square, both in fp32. That is 16 bytes per parameter before the first activation is allocated. A model with one billion parameters therefore needs 16 GB of persistent state, which is why a 24 GB card starts to feel small a little past a few hundred million parameters. Gradients do not exist until backward runs, and `optimizer.zero_grad(set_to_none=True)` frees them after each step instead of zero filling them. That does not lower the peak, since backward reallocates them, but the footprint between steps is lower and the first write into each buffer is a plain write rather than a read, add and write. On top of the persistent state there is workspace that no formula predicts well. The CUDA context takes several hundred megabytes that never appear in the allocator counters. cuDNN reserves workspace for convolution algorithms, and `torch.backends.cudnn.benchmark = True` will happily pick the fastest algorithm even when it wants a large scratch buffer. The caching allocator holds on to freed blocks for reuse, so `torch.cuda.memory_reserved()` runs ahead of `torch.cuda.memory_allocated()`, and fragmentation can leave the pool unable to serve one large request when the total free space would cover it; `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` fixes that more often than any other single change. ##### Activations, the part that scales with the batch Activations are every intermediate tensor that autograd keeps because the backward pass needs it. The output of each convolution is kept for that convolution's weight gradient. The input to each normalisation layer is kept. For a convolutional segmentation network at high resolution this is the dominant term by a wide margin: a handful of 1024 pixel crops through a backbone with dozens of stages produces activations measured in gigabytes, against a backbone whose weights fit in a few hundred megabytes. For transformers the published approximation is that each layer keeps roughly `s * b * h * (34 + 5 * a * s / h)` bytes in mixed precision without a fused attention kernel, where s is sequence length, b micro batch, h hidden size and a the head count; the second term is the score matrix that flash attention removes. The point is that activations scale with batch and sequence, not with parameter count, which is why halving the batch rescues a run that is slightly over and does nothing for one that is hopelessly over. The peak usually lands at the boundary between forward and backward, when every activation is alive and the first gradient buffers are being allocated; from there the backward frees activations layer by layer as it consumes them. Here is the estimate we run before choosing a configuration; it is deliberately rough about activations, which should be measured rather than predicted. ```python import torch _BYTES = {torch.float32: 4, torch.bfloat16: 2, torch.float16: 2} def training_memory_estimate(model, activation_bytes, *, optimizer="adamw", master_dtype=torch.float32, autocast_dtype=None, kept_fraction=1.0, workspace_bytes=1.5e9): """Rough peak in GiB. activation_bytes is the measured forward footprint of one micro batch at the chosen precision; kept_fraction is the share still stored once checkpointing discards the rest.""" n = sum(p.numel() for p in model.parameters() if p.requires_grad) w = _BYTES[master_dtype] moments = {"sgd": 0, "sgd_momentum": 1, "adamw": 2}[optimizer] parts = { "weights": n * w, "gradients": n * w, "optimiser_state": n * 4 * moments, # torch.optim keeps moments in fp32 "autocast_cache": n * _BYTES[autocast_dtype] if autocast_dtype else 0, "activations": activation_bytes * kept_fraction, "workspace": workspace_bytes, } parts["total"] = sum(parts.values()) return {k: v / 2**30 for k, v in parts.items()} ``` It says nothing about fragmentation or the CUDA context, which is why we treat its total as a floor and keep a margin of a couple of gigabytes. ##### Mixed precision: what bf16 and fp16 actually change `torch.autocast` does not convert the model. The weights stay in fp32, the optimiser state stays in fp32, and gradients accumulate into fp32 `.grad` buffers. What changes is that matrix multiplications and convolutions run on half precision copies of their inputs, and their outputs, which are the activations, are stored in half precision. So mixed precision roughly halves the activation term and leaves the 16P persistent term alone. There is one hidden cost: inside an autocast region PyTorch caches the half precision copy of each weight so it is not recast for every op, which is a transient 2P during the forward, released when the context exits. bf16 and fp16 are not interchangeable. fp16 spends 5 bits on the exponent and 10 on the mantissa, so it is precise but narrow: the smallest normal value is about 6e-5 and the largest is 65504. Gradients in a deep network routinely sit below 6e-5, so they underflow to zero and the run quietly stops learning without an error. Loss scaling is the fix: multiply the loss by a large factor S before backward so every gradient is scaled by S into the representable range, then divide by S before the optimiser step. `GradScaler` manages S dynamically. It starts high, checks the unscaled gradients for inf or nan after every backward, skips the step and halves S when it finds one, and doubles S after a run of clean steps. Skipped steps are normal early in training and should be logged, not hidden. bf16 spends 8 bits on the exponent, the same as fp32, and only 7 on the mantissa. Its range is that of fp32, so nothing underflows and no scaler is needed. Its precision is coarse, about three significant decimal digits, which is fine for activations and gradients and not fine for the weight update itself, which is why the master weights stay in fp32 and the optimiser step happens there. On Ampere and newer cards we use bf16 without a scaler, and the training step below keeps the scaler in but disabled so the same code runs with fp16 on an older card. ```figure gpu-memory-training-budget/memory-bar The same model at a fixed micro batch under fp32, bf16 autocast, and bf16 with checkpointing: the persistent state does not move while the activation term collapses. ``` ##### Gradient checkpointing and the recompute trade Checkpointing changes what autograd keeps. Wrap a block in `torch.utils.checkpoint.checkpoint` and the forward pass stores only the block's input, discarding everything computed inside it. When backward reaches the block it runs the block's forward again from the saved input to rebuild the intermediates, then backpropagates through them. The activation memory for a checkpointed block drops from everything inside it to its input tensor plus, transiently, one block's worth of internals during its own backward. The cost is compute: each checkpointed block runs its forward twice. The forward is roughly a third of a training step, so checkpointing every block costs around a third more wall time. Checkpoint the blocks with the largest activation to compute ratio, which for a segmentation backbone means the high resolution early stages rather than the small late ones. Use `use_reentrant=False`; the older re-entrant implementation cannot handle blocks whose inputs do not require gradients and composes badly with `torch.compile`. ```figure gpu-memory-training-budget/checkpoint-recompute Without checkpointing the forward keeps every intermediate for the backward; with it only block inputs are kept and each block's internals are rebuilt as the backward reaches it. ``` Two things to check afterwards. Random ops such as dropout must produce the same mask on recompute, which the non-reentrant implementation handles by preserving RNG state. And batch normalisation in training mode updates its running statistics on every forward, so a checkpointed block updates them twice per step, a slightly faster moving average that is harmless but worth knowing when comparing runs. ##### Gradient accumulation is not a bigger batch When the micro batch that fits is smaller than the batch you want, accumulate: run N micro batches, call backward on each with the loss divided by N, and step the optimiser once. The gradient buffers sum across the N backward calls, so the optimiser sees the mean gradient over N times the micro batch, which is what one large batch would have produced for any loss that is a mean over samples. Memory is one micro batch of activations plus the persistent state, since the accumulated gradients live in the same 4P buffer that was going to exist anyway; the one difference is that `set_to_none` cannot free them between micro batches, so the resident footprint stays higher for the whole group. The caveat is batch normalisation. BN computes its mean and variance over the tensor it is given, which is the micro batch, not the accumulated batch. With a micro batch of two, every BN layer normalises with statistics from two samples, and the running statistics used at inference average those noisy estimates. Accumulation does nothing to fix that, because the normalisation happens inside the forward pass before there is any gradient to accumulate. The honest options are to replace BN with GroupNorm or LayerNorm, which do not depend on the batch dimension, to freeze BN statistics from a pretrained backbone, or to accept the noise and watch the validation curve. ```figure gpu-memory-training-budget/accumulation-loop Each micro batch adds its scaled gradient into the same buffer; the optimiser steps and the buffer is released only after N micro batches. ``` ##### Memory format and fused optimisers `channels_last` stores a 4D tensor as NHWC instead of NCHW. It saves no memory, but with mixed precision on tensor cores cuDNN can run convolutions directly in that layout instead of inserting transposes around every op. Convert the model with `model.to(memory_format=torch.channels_last)` and every input with `x.contiguous(memory_format=torch.channels_last)`. The mistake to avoid is mixing formats: an NCHW input into an NHWC model makes PyTorch insert a copy at the boundary, a transient tensor's worth of memory and the time to write it. Check `x.is_contiguous(memory_format=torch.channels_last)` on the first batch. The optimiser step has two implementations on CUDA. The default `foreach` one groups parameters into lists and runs multi-tensor kernels, allocating intermediates while it does. `fused=True` runs the whole update in one kernel per parameter group with no intermediates, which is faster and removes the transient spike during the step. Since PyTorch 2.0 the fused AdamW also accepts the scaler's inverse scale and inf flag directly, so `GradScaler.step` does not need a separate unscale pass. ##### The training step This is the step we use, with all of the above in one place. ```python import torch import torch.nn as nn import torch.nn.functional as F from torch.utils.checkpoint import checkpoint class Stage(nn.Module): def __init__(self, block: nn.Module, recompute: bool): super().__init__() self.block, self.recompute = block, recompute def forward(self, x): if self.recompute and self.training and x.requires_grad: return checkpoint(self.block, x, use_reentrant=False) return self.block(x) def train_epoch(model, loader, opt, scaler, *, accum: int, dtype: torch.dtype, max_norm=1.0): model.train() opt.zero_grad(set_to_none=True) steps = len(loader) for i, (x, y) in enumerate(loader, start=1): x = x.cuda(non_blocking=True).contiguous(memory_format=torch.channels_last) y = y.cuda(non_blocking=True) with torch.autocast("cuda", dtype=dtype): logits = model(x) loss = F.cross_entropy(logits.float(), y, ignore_index=255) scaler.scale(loss / accum).backward() if i % accum == 0 or i == steps: scaler.unscale_(opt) nn.utils.clip_grad_norm_(model.parameters(), max_norm) scaler.step(opt) scaler.update() opt.zero_grad(set_to_none=True) model = build_model().cuda().to(memory_format=torch.channels_last) opt = torch.optim.AdamW(model.parameters(), lr=3e-4, weight_decay=0.05, fused=True) dtype = torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float16 scaler = torch.amp.GradScaler("cuda", enabled=(dtype == torch.float16)) ``` The loss is computed outside autocast in fp32; cross-entropy over half precision logits is the one place we have seen bf16's coarse mantissa show up as a noisier loss curve. Clipping happens after `unscale_`, because clipping scaled gradients would clip against a threshold that moves with S. The final partial group at the end of an epoch steps with a slightly smaller effective batch rather than being dropped; leaving a half accumulated gradient in the buffer for the next epoch to add to is the wrong thing to do. ##### Finding the peak with the memory snapshot When the estimate says a configuration fits and the card says otherwise, guessing is slower than looking. PyTorch records every allocation with a Python stack trace when asked: ```python torch.cuda.memory._record_memory_history(max_entries=200_000) train_epoch(model, short_loader, opt, scaler, accum=4, dtype=torch.bfloat16) torch.cuda.synchronize() torch.cuda.memory._dump_snapshot("step.pickle") torch.cuda.memory._record_memory_history(enabled=None) ``` Drop the pickle on the memory visualiser at pytorch.org/memory_viz and read the active memory timeline: each block is an allocation, its width is its lifetime, and hovering gives the stack that made it. Expect a staircase climbing through the forward, a plateau at the loss, and a staircase descending through the backward while the gradient buffers appear underneath; checkpointed blocks show as a small step in the forward and a short spike in the backward where the recompute happens. What we found the first time we did this on a segmentation run was none of those: the peak was a full resolution fp32 copy of the logits produced by an interpolation inside the loss function, twice the size of any backbone activation, alive for the whole backward because the loss held a reference to it. It was a one line fix, and no amount of batch size fiddling would have found it. ### Designing MCP Tool Servers an Agent Can Actually Use - Tags: MCP, Agents, TypeScript, LLM - Summary: What we changed in a Model Context Protocol tool server after the first week of real agent traffic: schemas as the real prompt, idempotency keys, budgets, an error taxonomy the model can act on, permission tiers, audit rows, and replay tests built from recorded traces. #### Designing MCP Tool Servers an Agent Can Actually Use The agentic workflows project put a multi-step LLM agent in front of a set of internal APIs and databases through a Model Context Protocol server. The first version worked in demos and came apart in the first week of real use, and almost none of the failures were in the model. They were in the tool surface: what the agent could see, what it was allowed to ask for, and what came back when something went wrong. This note is what we changed, in the order we learned it. ##### The tool schema is the real prompt From the model's side, an MCP server is three things: a `tools/list` response, a `tools/call` request, and whatever content comes back. The model never reads handler code. It reads names, descriptions, and a JSON schema, and it decides what to do from those alone. We spent the first two weeks polishing the system prompt and got less from that than from a single afternoon rewriting tool descriptions. This is the `create_refund` entry exactly as the agent sees it after `tools/list`: ```json { "name": "create_refund", "description": "Create one refund for one order. Call lookup_order first and use its refundableMinor value; amounts above it are rejected. Amounts are integers in minor units of the order currency. Generate a fresh idempotencyKey for each intended refund and reuse the same key when retrying the same refund.", "inputSchema": { "type": "object", "properties": { "orderId": { "type": "string", "pattern": "^ord_[a-z0-9]{12}$", "description": "From lookup_order" }, "amountMinor": { "type": "integer", "minimum": 1, "description": "At most refundableMinor" }, "reason": { "type": "string", "enum": ["damaged", "not_received", "duplicate_charge", "customer_request"] }, "idempotencyKey": { "type": "string", "format": "uuid", "description": "New UUID per refund; reuse on retry" } }, "required": ["orderId", "amountMinor", "reason", "idempotencyKey"], "additionalProperties": false }, "annotations": { "readOnlyHint": false, "destructiveHint": false, "idempotentHint": true } } ``` Three rules produced most of the improvement. Names are `verb_noun` and the verbs are the same everywhere: `lookup_`, `list_`, `create_`, `start_`, `get_`. Argument shapes are as narrow as the domain allows: money is an integer in minor units, never a float; dates are ISO 8601 strings with a pattern; identifiers carry a regex so a hallucinated id fails validation before it reaches a database. And anything that used to be free text became an enum. The `reason` field is the clearest example. As a string it came back with a different phrasing on nearly every call and the reporting team could not group them. As a four-value enum the model picks the right one and downstream code can switch on it. We register tools with the MCP TypeScript SDK and Zod, and the JSON above is generated from the Zod shape, so the schema the model reads and the validation the handler runs can never drift apart: ```ts import { McpServer } from '@modelcontextprotocol/sdk/server/mcp.js'; import { z } from 'zod'; const server = new McpServer({ name: 'ops-tools', version: '3.2.0' }); server.registerTool( 'create_refund', { title: 'Create refund', description: CREATE_REFUND_DESCRIPTION, inputSchema: { orderId: z.string().regex(/^ord_[a-z0-9]{12}$/).describe('From lookup_order'), amountMinor: z.number().int().positive().describe('At most refundableMinor'), reason: z.enum(['damaged', 'not_received', 'duplicate_charge', 'customer_request']), idempotencyKey: z.string().uuid().describe('New UUID per refund; reuse on retry'), }, annotations: { readOnlyHint: false, destructiveHint: false, idempotentHint: true }, }, (args, extra) => withPolicy('create_refund', args, extra, createRefund), ); ``` `withPolicy` is the wrapper that runs the permission tier, the budget check, the idempotency claim and the audit row before the handler is reached. Every tool goes through it; no handler implements those concerns itself. ##### Small orthogonal tools over one mega tool The first design had a single `orders` tool with an `action` argument and a union of argument shapes. It looked tidy in the code and it was the worst thing in the server. The model mixed fields between actions, the description grew until the useful sentences were buried, and there was no way to give `lookup` and `refund` different permission tiers because they were the same tool. We split it into one tool per action, each with a schema small enough to be honoured in full. The tool list got longer, which costs context on every turn, and that was worth it. Each tool can now carry its own annotations, its own tier, its own budget, and its own example in the description. The rule we settled on: if two operations need different permissions, different retry semantics, or different result shapes, they are different tools. ```figure mcp-tool-server-design/agent-loop The agent reads the tool list, emits a call, and reads a structured result; every call passes through validation, permission, budget and audit stages inside the server before a handler runs. ``` ##### Idempotency keys and safe retries Agents retry more than any human client. The transport retries on timeouts, the model retries when a result looks wrong, and an operator will happily re-run a whole conversation from a checkpoint. Any tool with a side effect has to be safe to call twice with the same intent, and "the same intent" has to be something the server can recognise. Every non-read tool takes an `idempotencyKey`. The description tells the model to mint a new UUID per intended action and to reuse it on retry, and in practice the model follows that instruction reliably once it is stated in the schema rather than the system prompt. Keys are scoped to the session. The handler claims the key in Postgres before doing anything, using a unique index so the claim is atomic, and stores a fingerprint of the arguments so a reused key with different arguments is rejected instead of silently replayed. ```ts import { createHash } from 'node:crypto'; import type { CallToolResult } from '@modelcontextprotocol/sdk/types.js'; interface SessionContext { sessionId: string; actor: string } interface ClaimRow { inserted: boolean; state: 'in_progress' | 'done'; fingerprint: string; result: Record | null } const ok = (data: Record, summary: string): CallToolResult => ({ content: [{ type: 'text', text: summary }], structuredContent: data, }); const fail = (error: ToolError): CallToolResult => ({ isError: true, content: [{ type: 'text', text: `${error.code}: ${error.message}` }], structuredContent: { error }, }); export async function createRefund(args: CreateRefundArgs, ctx: SessionContext): Promise { const fingerprint = createHash('sha256') .update(JSON.stringify([args.orderId, args.amountMinor, args.reason])) .digest('hex'); // xmax = 0 is true only for the row this statement inserted, so one round trip tells us claim vs replay. const claim = await db.one( `INSERT INTO tool_claims (session_id, key, tool, fingerprint, state) VALUES ($1, $2, 'create_refund', $3, 'in_progress') ON CONFLICT (session_id, key) DO UPDATE SET session_id = EXCLUDED.session_id RETURNING (xmax = 0) AS inserted, state, fingerprint, result`, [ctx.sessionId, args.idempotencyKey, fingerprint], ); if (!claim.inserted) { if (claim.fingerprint !== fingerprint) { return fail({ code: 'CONFLICT', retryable: false, message: 'This idempotencyKey was already used with different arguments. Use a new key for a new refund.' }); } if (claim.state === 'in_progress') { return fail({ code: 'IN_PROGRESS', retryable: true, retryAfterMs: 2000, message: 'A previous attempt with this key is still running. Call again with the same key.' }); } return ok({ ...claim.result, replayed: true }, 'This refund was already created; returning the stored result.'); } try { const refund = await ordersApi.createRefund({ ...args, requestId: args.idempotencyKey }); const result = { refundId: refund.id, status: refund.status, amountMinor: refund.amountMinor }; await db.none(`UPDATE tool_claims SET state = 'done', result = $3 WHERE session_id = $1 AND key = $2`, [ctx.sessionId, args.idempotencyKey, result]); return ok(result, `Refund ${refund.id} created for ${args.amountMinor} minor units.`); } catch { await db.none(`DELETE FROM tool_claims WHERE session_id = $1 AND key = $2 AND state = 'in_progress'`, [ctx.sessionId, args.idempotencyKey]); return fail({ code: 'UPSTREAM_UNAVAILABLE', retryable: true, retryAfterMs: 5000, message: 'The orders API did not respond. Retry with the same idempotencyKey.' }); } } ``` The key is also forwarded to the upstream API as its request id, so deduplication extends end to end. Where an upstream had no idempotency support of its own we did not delete the claim on failure; we marked it unknown and made the next call with that key run a lookup before deciding whether to execute. That is slower and it is the only honest option when you cannot tell whether the side effect happened. ```figure mcp-tool-server-design/idempotent-retry A first attempt succeeds upstream but its response is lost; the retry carries the same key, hits the stored claim, and is answered from the store without a second upstream call. ``` ##### Budgets and rate limits per session A confused loop will call the same tool hundreds of times, and each call costs money on the model side and load on the internal system. Every session gets a budget: total calls, calls per tool, wall clock, and domain budgets such as the total refund value a single session may create. The counters live in Redis keyed by session and tool, checked inside `withPolicy` before the handler. The distinction that mattered was between a budget and a rate limit. `RATE_LIMITED` comes back with `retryable: true` and a `retryAfterMs`, and the model waits and continues. `BUDGET_EXHAUSTED` comes back with `retryable: false` and a message that tells the model to summarise what it has done and stop. Before we separated the two, both were a generic "try again later", and the model did exactly that until the session was killed. ##### Structured results and an error taxonomy the model can act on Every result has two halves: a short text summary for the transcript and a `structuredContent` object for the next step. Errors set `isError` and carry a typed object rather than a message string: ```ts type ToolErrorCode = | 'INVALID_ARGUMENT' | 'NOT_FOUND' | 'CONFLICT' | 'IN_PROGRESS' | 'PERMISSION_REQUIRED' | 'BUDGET_EXHAUSTED' | 'RATE_LIMITED' | 'UPSTREAM_UNAVAILABLE'; interface ToolError { code: ToolErrorCode; retryable: boolean; message: string; // what to do next, written for the model retryAfterMs?: number; details?: Record; } ``` The message says what to do next, not what went wrong internally. Stack traces are worse than useless: they are long, they consume context, and the model tries to fix them. `NOT_FOUND` says which id was not found and which lookup tool to call. `INVALID_ARGUMENT` names the field and the constraint, which the schema validator gives us for free. A small closed set of codes also means the audit log can be grouped by code, which is how we found that most `CONFLICT` errors came from one description that had not told the model to mint a fresh key per action. ##### Pagination and cursors No list tool returns everything. `list_` tools take `cursor` and `limit`, enforce a maximum limit chosen for the model's context rather than for the database, and return `items`, `nextCursor` and `hasMore`. The cursor is an opaque, signed, base64 encoding of a keyset position, so the model cannot invent one or increment it, and a stale cursor fails validation with `INVALID_ARGUMENT` rather than returning a random page. We do not return totals unless the count is cheap; a count over a large table for every page is a good way to spend the latency budget on a number the model rarely uses. ##### Long running tasks with progress Report exports took minutes, which is longer than any tool call should block. The model-facing contract is a pair of tools: `start_export` validates its arguments, enqueues the job and returns `{ jobId, state: 'queued', expectedSeconds }`, and `get_job` returns the state, a progress fraction, and, once finished, the result or a resource link. MCP also has progress notifications tied to a `progressToken`, and we emit them for the human review UI, but we did not make the model depend on them because not every client surfaces notifications to the model. The `expectedSeconds` field matters more than it looks: without it the model polled in a tight loop and burned its call budget waiting. ##### Permission tiers for side effects Tools are registered with a tier: `read` runs automatically, `write` passes through a confirmation gate, and `destructive` requires a human approval recorded out of band. The tier is declared once at registration and enforced in `withPolicy`, never inside a handler, so a new tool cannot forget to check. The annotations in the schema mirror the tier so clients that render tool lists can show it. The confirmation flow reuses the idempotency machinery. A `write` call that needs confirmation returns `PERMISSION_REQUIRED` with a `confirmationId` and a plain-language description of the side effect. The operator approves or rejects it in the review UI. The model then calls the same tool again with the same `idempotencyKey` and the `confirmationId`, the server matches the claim, and the handler runs. A rejected confirmation leaves the claim in a terminal state so the model cannot re-ask for the same action with the same key. ```figure mcp-tool-server-design/permission-gate A write-tier call is held at the gate; the server returns a confirmation id, an operator approves, and the agent retries with the same idempotency key so the approved action runs exactly once. ``` ##### Audit logging Every `tools/call` produces an audit row: session, actor, tool, arguments with fields marked sensitive in the schema redacted, permission decision, idempotency outcome, error code, latency, and the upstream request ids. The row is written before the handler runs and updated when it returns, so a process that dies mid-handler leaves a visible unfinished row rather than nothing. The table is append-only for everything except that completion update. We built the audit log for compliance and it became the primary debugging tool. Almost every bug report started as a session id, and the sequence of rows for that session usually explained the behaviour without a single log line from the model side. It is also where the test traces come from. ##### Testing a tool server with recorded agent traces Unit tests of individual handlers were never the problem. The bugs that hurt were sequences: the model calls `lookup_order`, gets a `NOT_FOUND`, calls `list_orders`, picks a stale id from the wrong page, and calls `create_refund` with it. Those only show up when the whole server runs against a realistic sequence of calls. A trace is exported from the audit log: the ordered tool calls with their arguments, the results the server produced, and the recorded upstream responses for that session. A replay test feeds the sequence back through a real server instance with a fixed clock and a stubbed upstream that answers from the recording, then asserts the results and, just as important, the exact list of upstream calls that were made. ```ts import { test } from 'node:test'; import assert from 'node:assert/strict'; import { readFileSync } from 'node:fs'; import { createTestHarness, errorCode } from './harness'; interface TraceStep { tool: string; args: Record; expect: { isError: boolean; code?: string; structured?: Record }; } interface Trace { sessionId: string; startedAt: string; upstream: Record; steps: TraceStep[]; upstreamCalls: string[]; } const trace: Trace = JSON.parse(readFileSync(new URL('./traces/refund-after-timeout.json', import.meta.url), 'utf8')); test('refund-after-timeout replays without a second upstream refund', async () => { const h = createTestHarness({ now: new Date(trace.startedAt), upstream: trace.upstream }); for (const [i, step] of trace.steps.entries()) { const res = await h.callTool(step.tool, step.args, { sessionId: trace.sessionId, actor: 'replay' }); assert.equal(Boolean(res.isError), step.expect.isError, `step ${i}: error flag`); if (step.expect.code) assert.equal(errorCode(res), step.expect.code, `step ${i}: error code`); if (step.expect.structured) assert.deepEqual(res.structuredContent, step.expect.structured, `step ${i}: payload`); } assert.deepEqual(h.upstreamCalls.map((c) => c.name), trace.upstreamCalls); }); ``` The replay suite is what makes schema changes safe. When a description or a shape changes, recorded calls that no longer validate fail loudly, and that is the point: it is a deliberate break, and the trace is regenerated by running the model against the new server and reviewing the new sequence before it is committed. Traces that encode a bug we fixed are kept as regression tests with the expected results updated, so the sequence that once double-refunded an order now asserts that the second attempt is a replay. ##### What is still open Descriptions are code, and we still review them less carefully than code. The change that most improved the agent's behaviour was a sentence in a description, and the change that most degraded it was also a sentence in a description, both of which went through review as trivial diffs. The next step is to run the replay suite on description changes with the live model, not only the recorded calls, so a reworded tool gets exercised before it ships. ### Row Level Security Done Properly: Multi Tenant Isolation in Postgres and Supabase - Tags: PostgreSQL, Supabase, Security, Backend - Summary: Why tenant filters in application code leak, how USING and WITH CHECK are evaluated, where the tenant id comes from, the per row function and missing index traps and the InitPlan fix, a policy test script, service role discipline and migrations that do not lock everyone out. #### Row Level Security Done Properly: Multi Tenant Isolation in Postgres and Supabase Every multi tenant application starts with the same sentence in a design document: every query filters by tenant. Every one of them eventually ships a query that does not. The filter lives in application code, and application code has an ORM that lazily loads a relation, a reporting endpoint written in a hurry, an admin script that someone ran with production credentials, and a new engineer who did not know the rule. Row level security moves the rule into the database, where it applies to every statement from every client whether or not the author remembered it. It is the right tool, and it is also easy to deploy in a way that is slow, incomplete, or silently bypassed. This note is how we set it up, the traps we hit, and how we test it. ##### Why the filter cannot live in the application The problem with `WHERE tenant_id = $1` in the application is not that it is hard to write. It is that it has to be written everywhere, and the set of places grows faster than anyone reviews it. A join that starts from a table with a tenant filter and reaches one without it leaks. A subquery in a `count` used for pagination leaks. A raw SQL string in a migration backfill leaks. The failure is quiet: no error, just rows from another customer in a response, discovered by that customer. Application filters remain useful as an optimisation and as a first line, but they cannot be the boundary. The boundary has to be enforced by the component that sees every query, and that is the database. ##### How a policy is evaluated Enabling RLS on a table is one statement, and it changes the semantics of every query against that table for every role except the owner and superusers. Once enabled, a table with no policies returns nothing and accepts nothing; RLS is default deny. Policies then grant access back, each with two expressions. `USING` is evaluated against rows that already exist and decides whether a `SELECT` returns them and whether an `UPDATE` or `DELETE` can target them. `WITH CHECK` is evaluated against rows being written and decides whether an `INSERT` or the result of an `UPDATE` is allowed to exist. If a policy for `INSERT` or `UPDATE` omits `WITH CHECK`, Postgres reuses the `USING` expression for it, which is convenient right up to the moment a policy grants broad read access and thereby grants the same breadth of write access. Policies are permissive by default and combine with `OR`: a row passes if any permissive policy for that command and role accepts it. Restrictive policies, declared with `AS RESTRICTIVE`, combine with `AND` on top: every restrictive policy must also accept the row. The rule is that at least one permissive policy must pass and no restrictive policy may fail. Restrictive policies are the right tool for a tenant boundary that has to hold regardless of what feature specific policies are added later; a permissive policy for a new feature can widen access within a tenant but can never cross a restrictive tenant check. Policies filter rows; they do not replace grants. A role still needs `SELECT`, `INSERT`, `UPDATE` or `DELETE` on the table, and RLS decides which rows those privileges reach. Two more details determine whether the boundary is real. The table owner bypasses RLS unless `ALTER TABLE ... FORCE ROW LEVEL SECURITY` is set, and in a Supabase project the owner is typically `postgres`, which is what migrations and many server side scripts connect as. Roles with the `BYPASSRLS` attribute, which includes the Supabase `service_role`, skip every policy. ##### Where the tenant comes from A policy needs to know who is asking. In Supabase the API layer verifies the JWT and sets its claims as a transaction local configuration parameter, which `auth.uid()` reads. Under the hood it is `current_setting('request.jwt.claims', true)::jsonb ->> 'sub'`, and any custom claim added to the token at sign in, a `tenant_id` for example, is available the same way. For a self hosted Postgres behind our own API we do the same thing by hand: at the start of each transaction the API runs `select set_config('app.user_id', $1, true)` and policies read `current_setting('app.user_id', true)::uuid`. The third argument, `is_local`, scopes the value to the transaction so it cannot outlive it. That matters when a connection pooler in transaction mode hands the same backend to a different request a millisecond later; a session level `SET` would carry one tenant's identity into another tenant's query. The design we settled on is a memberships table linking users to tenants, and a helper function that returns the tenant ids the current user belongs to. ```sql create schema if not exists app; create table public.tenants ( id uuid primary key default gen_random_uuid(), name text not null ); create table public.memberships ( tenant_id uuid not null references public.tenants (id) on delete cascade, user_id uuid not null, role text not null default 'member' check (role in ('member', 'owner')), primary key (tenant_id, user_id) ); create index memberships_user_idx on public.memberships (user_id); create table public.projects ( id uuid primary key default gen_random_uuid(), tenant_id uuid not null references public.tenants (id), name text not null, created_by uuid not null, created_at timestamptz not null default now() ); create index projects_tenant_idx on public.projects (tenant_id, created_at desc); -- Tenants the current user belongs to, optionally with a required role. create or replace function app.my_tenant_ids(required_role text default null) returns uuid[] language sql stable security definer set search_path = '' as $$ select coalesce(array_agg(m.tenant_id), '{}') from public.memberships m where m.user_id = (select auth.uid()) and (required_role is null or m.role = required_role); $$; revoke all on function app.my_tenant_ids(text) from public; grant execute on function app.my_tenant_ids(text) to authenticated; alter table public.memberships enable row level security; alter table public.projects enable row level security; alter table public.projects force row level security; create policy memberships_self on public.memberships for select to authenticated using (user_id = (select auth.uid())); create policy projects_tenant_read on public.projects for select to authenticated using (tenant_id = any ((select app.my_tenant_ids()))); create policy projects_tenant_insert on public.projects for insert to authenticated with check ( tenant_id = any ((select app.my_tenant_ids())) and created_by = (select auth.uid()) ); create policy projects_tenant_update on public.projects for update to authenticated using (tenant_id = any ((select app.my_tenant_ids()))) with check (tenant_id = any ((select app.my_tenant_ids()))); create policy projects_owner_delete on public.projects for delete to authenticated using (tenant_id = any ((select app.my_tenant_ids('owner')))); ``` The helper is `security definer` so that it reads memberships as its owner, which means the memberships table's own policy is not evaluated inside every projects policy, and `stable` so the planner knows it returns the same value for every row of one statement. `set search_path = ''` with schema qualified names inside is not optional on a definer function: without it a caller who can create objects in a schema earlier on the path can substitute their own `memberships`. The `update` policy has both clauses because a user must be able to see the row they are changing and the changed row must still belong to one of their tenants; without the `WITH CHECK`, an update that moves `tenant_id` to another customer's id would be accepted. ```figure postgres-rls-multi-tenant/policy-gate Every row the scan produces passes through the policy's USING expression; rows from tenants the caller does not belong to are dropped before the query sees them. ``` ##### The performance traps A policy expression is appended to the query as a predicate, and it is evaluated the way any predicate is: potentially once per row. That is where RLS deployments go wrong. The first trap is calling a function per row. `auth.uid()` looks like a constant but it is a function that parses a JSON setting, and written directly as `user_id = auth.uid()` the executor may call it for every row the scan visits. Wrapping it in a scalar subquery, `(select auth.uid())`, turns it into an `InitPlan`: an uncorrelated subquery the planner evaluates once and reuses as a parameter. The same applies to the helper. `tenant_id = any ((select app.my_tenant_ids()))` runs the membership lookup once per statement and compares every row against the resulting array, and because the comparison is a plain equality against a parameter, the index on `tenant_id` can serve it. The second trap is a correlated subquery. A policy written as `exists (select 1 from memberships m where m.tenant_id = projects.tenant_id and m.user_id = auth.uid())` reads naturally, and it forces a subplan that references the outer row, so it runs once per row of `projects`; inside each of those runs the memberships table's own RLS policy is evaluated and `auth.uid()` is called again. On a few thousand rows nobody notices. On a table with millions of rows shared across tenants it turns a millisecond lookup into a sequential scan with a nested loop. Here is the difference in plan shape, with costs turned off so the structure is what you read. ```sql -- Correlated policy: the membership check runs per row. explain (costs off) select id, name from public.projects where created_at > timestamptz '2026-08-29'; -- Seq Scan on projects -- Filter: ((created_at > '2026-08-29 00:00:00+00'::timestamptz) AND (SubPlan 1)) -- SubPlan 1 -- -> Index Only Scan using memberships_pkey on memberships m -- Index Cond: ((tenant_id = projects.tenant_id) AND (user_id = auth.uid())) -- InitPlan policy: one membership lookup, then an index scan on the tenant column. explain (costs off) select id, name from public.projects where created_at > timestamptz '2026-08-29'; -- Index Scan using projects_tenant_idx on projects -- Index Cond: ((tenant_id = ANY ($0)) AND (created_at > '2026-08-29 00:00:00+00'::timestamptz)) -- InitPlan 1 (returns $0) -- -> Result ``` The second plan is what a tenant scoped query should look like: the identity is resolved once, and the composite index on `(tenant_id, created_at desc)` answers the policy predicate and the application's own predicate in one index condition. ```figure postgres-rls-multi-tenant/initplan-vs-per-row A correlated policy calls into the membership check for each row the scan visits; the InitPlan form resolves the caller's tenants once and reuses the result as a parameter. ``` The third trap is the missing index. Once RLS is on, every query against the table carries the `tenant_id` predicate whether the application wrote one or not, and a table without an index leading on `tenant_id` will scan. Indexes that served the application's own filters often need `tenant_id` prepended so both predicates land in one index condition. The fourth is subtler. To stop a malicious user extracting data through error messages, the planner refuses to evaluate a user supplied predicate before the policy predicate unless the functions and operators involved are marked `leakproof`. Equality on integers, uuids and text is leakproof. `LIKE`, most casts and every user defined function are not. If the application's `WHERE` uses a non leakproof operator, that condition is applied after the policy check, and an index that would have served it alone may no longer be chosen. The fix is usually to make the policy cheap enough that it does not matter, or to add the tenant column to the index the application predicate wants. Views are the last trap and the one most specific to Supabase. A view runs with the privileges of its owner by default, and if the owner is `postgres` the view reads the underlying tables with RLS bypassed. Since Postgres 15, `create view ... with (security_invoker = true)` makes the view evaluate policies as the caller, and we set it on every view that touches a tenant table. ##### Testing policies like code A policy is a piece of authorisation logic and deserves a test that fails when it regresses. Postgres gives us everything needed: `SET ROLE` switches to the role the API uses, `set_config` plants the JWT claims the policy reads, and a transaction that ends in `ROLLBACK` leaves no trace. This is the script we run in CI against a database that has just had the migrations applied. ```sql begin; -- Fixture: two tenants, one user in each, one user in both. insert into public.tenants (id, name) values ('11111111-1111-1111-1111-111111111111', 'acme'), ('22222222-2222-2222-2222-222222222222', 'globex'); insert into public.memberships (tenant_id, user_id, role) values ('11111111-1111-1111-1111-111111111111', 'aaaaaaaa-0000-0000-0000-000000000001', 'owner'), ('22222222-2222-2222-2222-222222222222', 'aaaaaaaa-0000-0000-0000-000000000002', 'member'), ('11111111-1111-1111-1111-111111111111', 'aaaaaaaa-0000-0000-0000-000000000003', 'member'), ('22222222-2222-2222-2222-222222222222', 'aaaaaaaa-0000-0000-0000-000000000003', 'member'); insert into public.projects (tenant_id, name, created_by) values ('11111111-1111-1111-1111-111111111111', 'acme one', 'aaaaaaaa-0000-0000-0000-000000000001'), ('11111111-1111-1111-1111-111111111111', 'acme two', 'aaaaaaaa-0000-0000-0000-000000000001'), ('22222222-2222-2222-2222-222222222222', 'globex one', 'aaaaaaaa-0000-0000-0000-000000000002'); create function pg_temp.login(uid uuid) returns void language sql as $$ select set_config('request.jwt.claims', json_build_object('sub', uid, 'role', 'authenticated')::text, true); $$; create function pg_temp.assert_count(expected int, msg text) returns void language plpgsql as $$ declare n int; begin select count(*) into n from public.projects; if n <> expected then raise exception 'FAIL %: expected % rows, saw %', msg, expected, n; end if; raise notice 'ok: %', msg; end $$; set local role authenticated; select pg_temp.login('aaaaaaaa-0000-0000-0000-000000000001'); select pg_temp.assert_count(2, 'acme owner sees only acme projects'); select pg_temp.login('aaaaaaaa-0000-0000-0000-000000000003'); select pg_temp.assert_count(3, 'dual member sees both tenants'); select pg_temp.login('aaaaaaaa-0000-0000-0000-000000000002'); select pg_temp.assert_count(1, 'globex member sees one project'); -- Cross tenant insert must be rejected by WITH CHECK. do $$ begin insert into public.projects (tenant_id, name, created_by) values ('11111111-1111-1111-1111-111111111111', 'smuggled', 'aaaaaaaa-0000-0000-0000-000000000002'); raise exception 'FAIL: cross tenant insert was accepted'; exception when insufficient_privilege then raise notice 'ok: cross tenant insert rejected'; end $$; -- A member cannot delete; the statement succeeds and affects zero rows. delete from public.projects where name = 'globex one'; select pg_temp.assert_count(1, 'member delete affected no rows'); rollback; ``` The insert test is the one that matters most and the one most often missing. A `WITH CHECK` failure raises SQLSTATE 42501, `insufficient_privilege`, and the test asserts that specific exception rather than any error, so a typo in the fixture cannot pass as a policy rejection. The delete test shows the other side of RLS: a `DELETE` or `UPDATE` whose target rows fail `USING` does not error, it simply matches nothing, so the assertion has to be on the row count. ```figure postgres-rls-multi-tenant/with-check-write An insert carrying another tenant's id reaches the WITH CHECK expression and is rejected with 42501; the same insert for the caller's own tenant is written. ``` ##### Service role discipline Supabase's `service_role` key bypasses RLS entirely, and it exists because some code legitimately needs to: migrations, cron jobs that touch every tenant, support tooling. The discipline is simple to state and hard to keep. The key never reaches a browser, a mobile bundle or an edge function that handles user requests. Server code that acts on behalf of a user uses the user's JWT, so the same policies apply and the server cannot leak more than the user could see. Code that genuinely needs to cross tenants uses the service key through one module that logs what it did and why. On a self hosted database the equivalent is to create the application role with `NOBYPASSRLS`, grant it only the tables it needs, and keep a separate role with `BYPASSRLS` whose credentials are not in the application's environment at all. ##### Migrations that add policies safely Enabling RLS on a live table is default deny from the instant the statement commits. A migration that enables RLS in one step and adds policies in the next, with anything in between, is an outage for every tenant. The order we use, all in one transaction, is: backfill `tenant_id` so it can be declared `NOT NULL`, create the index that leads on it, create the policies, then enable and force RLS. `CREATE POLICY` on a table that does not yet have RLS enabled is allowed and does nothing until it is, so the policies are in place before the switch flips. `ALTER TABLE ... ENABLE ROW LEVEL SECURITY` takes an `ACCESS EXCLUSIVE` lock. The lock itself is held for microseconds, but it queues behind any long running query on the table and every query after it queues behind the lock, so we set `lock_timeout` to a couple of seconds and let the migration retry rather than stall the application. Policies cannot be created idempotently, so each migration drops its own policies with `DROP POLICY IF EXISTS` before creating them, which makes the migration re-runnable and makes the policy definition in the migration file the single source of truth. And before the migration ever runs against production, the test script above runs against a copy with the migration applied, inside a transaction, so that a wrong policy fails a build rather than a customer. ### Build the RAG Evaluation Harness Before You Touch the Prompt - Tags: RAG, Evaluation, LLM, Python - Summary: How the evaluation harness for an enterprise RAG pipeline is put together: a golden set from real questions, recall at k, MRR and nDCG with their blind spots, a calibrated LLM judge for faithfulness and citation precision, cost and latency budgets, and the CI gate that caught a change which raised recall while quietly breaking grounding. #### Build the RAG Evaluation Harness Before You Touch the Prompt The enterprise RAG pipeline answers questions over a private company knowledge base: structure aware chunking, embeddings in pgvector, hybrid dense plus keyword retrieval, a cross-encoder reranker, and answers grounded with inline citations. For the first few weeks we tuned the answer prompt against a handful of questions we happened to know. Every change looked like an improvement to the person who made it, and two people could not agree on whether the same change was better or worse. We stopped and built the evaluation harness, and everything that followed went through it. This is how it is put together and what it caught. ##### Why prompt tuning without retrieval metrics is guessing A RAG answer can fail in two separate places. Either the passage that contains the answer was not retrieved, or it was retrieved and the model did not use it properly. No prompt fixes the first failure. If you cannot see which of the two happened, you attribute retrieval misses to the prompt, and the instructions you add to compensate make the model more confident with the wrong context. That is a worse system that scores better on the three questions you tried it on. So the harness measures retrieval on its own first, then generation given the retrieved passages, and reports both. A change is allowed to move only one of them, and when it moves both, we want to know before it ships. ##### Building a golden set from real questions The golden set is the asset everything else depends on, and it has to come from real questions. Ours came from three sources: the pilot's query logs, a backlog of internal support tickets that had been answered by pointing at a document, and a list the domain experts kept of questions they were asked repeatedly. Questions we invented ourselves were kept in a separate smoke set and never used for decisions, because we found we invented questions that our own chunker happened to handle well. Each item holds the question, the set of relevant passages with a relevance grade, and a reference answer written by an expert with the citations they would expect. Two people label each item independently and disagreements are adjudicated rather than averaged. The set is stratified by document type and by how many passages the answer needs, because a set dominated by single-passage questions hides the multi-passage failures that matter most. One detail cost us a week: chunk ids are not stable. Change the chunker and every id changes, and a golden set keyed by chunk id becomes useless. Relevance is therefore recorded as document id plus character span, and it is mapped to chunk ids at evaluation time by span overlap against the index snapshot under test. The golden set is versioned alongside the index snapshot and the two are always reported together. ```figure evaluating-rag-before-tuning/harness-flow Golden questions flow through retrieval, reranking and generation; the relevant spans and reference answers feed the metrics and the judge, and every stage reports into one scoreboard. ``` ##### Retrieval metrics and what each hides We report three retrieval metrics and never one alone, because each hides something specific. Recall at k is the share of relevant passages that appear in the top k. It is the metric that tells you whether the generator even had a chance. It hides ordering: a relevant passage at rank 1 and at rank k count the same, and the generator does not treat them the same. MRR, the mean reciprocal rank of the first relevant passage, rewards getting one good passage to the top. It hides everything after the first hit. A question that needs three passages scores a perfect 1.0 if one of them is at rank 1 and the other two are missing. nDCG uses the relevance grades and discounts by position, so it is the fairest single number for ranking quality. It hides absolute coverage: a ranking can have excellent nDCG at k while still missing a passage the answer cannot do without. ```python from math import log2 from typing import Callable, Iterable, Mapping, Sequence def recall_at_k(ranked: Sequence[str], relevant: set[str], k: int) -> float: """Share of relevant chunk ids present in the top k.""" if not relevant: raise ValueError("golden item has no relevant chunks; fix the label, do not score it") hits = sum(1 for cid in ranked[:k] if cid in relevant) return hits / len(relevant) def mrr(ranked: Sequence[str], relevant: set[str]) -> float: """Reciprocal rank of the first relevant chunk, 0.0 if none is retrieved.""" for rank, cid in enumerate(ranked, start=1): if cid in relevant: return 1.0 / rank return 0.0 def ndcg_at_k(ranked: Sequence[str], grades: Mapping[str, int], k: int) -> float: """Graded, position-discounted gain normalised by the ideal ordering.""" dcg = sum(grades.get(cid, 0) / log2(rank + 1) for rank, cid in enumerate(ranked[:k], start=1)) ideal = sorted(grades.values(), reverse=True)[:k] idcg = sum(g / log2(rank + 1) for rank, g in enumerate(ideal, start=1)) return dcg / idcg if idcg else 0.0 def score_retrieval( golden: Iterable["GoldenItem"], retrieve: Callable[[str, int], list[str]], ks: tuple[int, ...] = (1, 3, 5, 8, 10), ) -> list[dict]: rows = [] for item in golden: ranked = retrieve(item.question, max(ks)) relevant = set(item.grades) # grades: chunk id -> 1 or 2, mapped from spans for this index snapshot row = {"id": item.id, "slice": item.slice, "mrr": mrr(ranked, relevant), "ndcg@10": ndcg_at_k(ranked, item.grades, 10)} row.update({f"recall@{k}": recall_at_k(ranked, relevant, k) for k in ks}) rows.append(row) return rows ``` The rows are aggregated per slice, not only overall. The overall recall moved very little on most changes; the per-slice numbers are where a chunker change showed up as a gain on prose documents and a loss on tables. ```figure evaluating-rag-before-tuning/recall-curve Recall at k steps up as k grows; the dashed line marks the k actually passed to the generator, and anything the curve gains to the right of it never reaches the model. ``` ##### Chunk level versus document level judgement Document-level recall is the number that looks best and means least. A long policy document can be "retrieved" because one of its chunks about a different section made the top k, and the chunk with the answer is nowhere. We track both. Document-level recall tells you whether the index and the embedding model know which document the question is about. Chunk-level recall, judged by span overlap, tells you whether the chunker cut the answer in a way that can be found. When document recall is high and chunk recall is low, the fix is in the chunker, not in the retriever, and we would have spent weeks on the wrong component without the split. ##### Answer faithfulness and citation precision with a calibrated judge Given the retrieved passages, the generation half is judged on two things. Faithfulness: is every claim in the answer supported by the passages it was given. Citation precision: does each citation marker actually support the sentence it is attached to. Both are judged by an LLM with a strict output schema, and the judge sees only the question, the passages and the answer. It does not see the reference answer, because faithfulness is about grounding, not correctness; a separate correctness judge uses the reference and is reported separately. ```python import json import anthropic JUDGE_MODEL = "claude-opus-5" JUDGE_PROMPT_VERSION = "faithfulness-v4" # bump on any change to JUDGE_SYSTEM or the schema FAITHFULNESS_SCHEMA = { "type": "object", "properties": { "claims": { "type": "array", "items": { "type": "object", "properties": { "text": {"type": "string"}, "verdict": {"type": "string", "enum": ["supported", "unsupported", "contradicted"]}, "supporting_passages": {"type": "array", "items": {"type": "string"}}, }, "required": ["text", "verdict", "supporting_passages"], "additionalProperties": False, }, }, "citation_checks": { "type": "array", "items": { "type": "object", "properties": { "citation": {"type": "string"}, "sentence": {"type": "string"}, "supports_sentence": {"type": "boolean"}, }, "required": ["citation", "sentence", "supports_sentence"], "additionalProperties": False, }, }, }, "required": ["claims", "citation_checks"], "additionalProperties": False, } JUDGE_SYSTEM = """You grade whether an answer is grounded in the passages it was given. Split the answer into atomic factual claims; ignore hedges and formatting. A claim is "supported" only if a passage states it or it follows directly from a passage. Background knowledge does not count. A claim is "contradicted" if a passage states the opposite. Otherwise it is "unsupported". For every citation marker such as [p12], decide whether that passage supports the sentence it is attached to. Do not judge whether the answer is correct or complete. Judge only whether it is grounded.""" def judge_faithfulness(client: anthropic.Anthropic, question: str, passages: list["Passage"], answer: str) -> dict: context = "\n\n".join(f"[{p.id}] {p.text}" for p in passages) response = client.messages.create( model=JUDGE_MODEL, max_tokens=4096, system=JUDGE_SYSTEM, messages=[{"role": "user", "content": f"Question:\n{question}\n\nPassages:\n{context}\n\nAnswer:\n{answer}"}], output_config={"format": {"type": "json_schema", "schema": FAITHFULNESS_SCHEMA}}, ) text = next(block.text for block in response.content if block.type == "text") data = json.loads(text) claims = data["claims"] checks = data["citation_checks"] return { "faithfulness": sum(c["verdict"] == "supported" for c in claims) / len(claims) if claims else 1.0, "citation_precision": sum(c["supports_sentence"] for c in checks) / len(checks) if checks else 0.0, "contradictions": sum(c["verdict"] == "contradicted" for c in claims), "judge": f"{JUDGE_MODEL}/{JUDGE_PROMPT_VERSION}", } ``` A judge is a model, so it is measured like one. A labelled subset of answers was graded claim by claim by two people, and the judge's verdicts were compared against those labels with agreement statistics until the prompt was stable, then the judge model and prompt version were frozen and stamped on every score row. Changing the judge is a change to the ruler, so it is done separately from changing the pipeline, and the old and new judge are run side by side on the same answers before the switch. We also learned to run the judge more than once on the same answer when a slice showed high variance, and to report the spread rather than a single number. ##### Regression gates in CI Every pull request that touches chunking, embeddings, retrieval, reranking, or the prompts runs the harness on a fixed subset of the golden set; the full set runs nightly. The gate compares against the baseline stored for the main branch and blocks the merge if recall at the served k drops beyond a tolerance, if faithfulness or citation precision drops at all beyond noise, or if cost per query or p95 latency exceed their budgets. The tolerance is the part people get wrong. A judge-based metric has run-to-run variance even on the same commit, so a tolerance chosen by taste either blocks harmless changes or lets real regressions through. We measured the variance by running the harness repeatedly on the same commit and set each gate's tolerance from the spread we observed. When the spread was too wide to be useful, the fix was in the judge, not in the tolerance. ```figure evaluating-rag-before-tuning/ci-gate A pull request runs the harness; recall and MRR rise, faithfulness falls below its floor, and the merge gate blocks the change even though the retrieval numbers improved. ``` ##### Cost and latency as first class metrics A change that widens the candidate pool or adds a reranking pass can be a pure improvement on recall and a regression in everything a user notices. So the harness records tokens per query split by retrieval context and generation, and wall time split by retrieval, rerank and generation, for every item, and reports the mean and the p95 next to the quality metrics. Widening the pool doubled reranker time in one experiment, and the recall it bought was in the part of the ranking the generator never saw. Cost and latency have gates like any other metric. Without them, quality metrics drift up and the bill and the response time drift up with them, and nobody notices until a user does. ##### A worked example: recall up, faithfulness down The change that justified the harness came from a reasonable idea. Increase the number of passages passed to the generator from five to eight, and add a query expansion step so that questions phrased differently from the documents still find their passages. On the retrieval side it did what it promised: recall at eight rose, and it rose most on the slice of questions phrased in everyday language rather than in the vocabulary of the documents. Faithfulness fell on the same run, and citation precision fell further. The reason took an afternoon of reading judged answers. The knowledge base held several versions of some documents, and the wider pool made it more likely that two versions of the same section appeared in the context together. The model blended statements from both, and its citations pointed at whichever version it had read last, which was often the superseded one. Every one of those answers read confidently, and none of them would have been caught by a demo. The fix was in retrieval, not in the prompt. The reranker now deduplicates by document version and caps the number of chunks from any single document, and each chunk carries its version and effective date in the metadata the generator can see. Recall at eight stayed where the change had put it, faithfulness recovered to the baseline, and citation precision ended above it because the version field gave the model a reason to cite the right copy. We also kept the failing run as a permanent regression item, so that combination of a wide pool and versioned duplicates can never come back unnoticed. ##### What the harness changed about how we work The practical effect is that arguments about the prompt ended. A proposal now arrives as a harness run, with the retrieval and generation numbers per slice, the cost and latency deltas, and a handful of judged answers where the numbers moved. Most proposals are still rejected, but they are rejected for a reason that is written down and can be checked, and the ones that ship are the ones that moved the right number without moving the wrong one. ### Streaming Model Output Without Leaking Resources: SSE, Backpressure and Cancellation - Tags: Next.js, Streaming, TypeScript, LLM - Summary: Zenith streams every reply through a Next.js route handler. This note covers the transport choice, propagating AbortSignal upstream, backpressure with pull, heartbeats, resuming with Last-Event-ID, and stopping generation the moment nobody is reading. #### Streaming Model Output Without Leaking Resources: SSE, Backpressure and Cancellation Zenith, the assistant widget on this site, answers through a Next.js route handler that forwards each request to a model on OpenRouter and streams the reply back token by token. The first version of that handler did the obvious thing: await the upstream stream, push every delta into a `ReadableStream`, return it. It worked in the browser and it was quietly wrong in three ways. It kept generating after the visitor closed the tab, it had no idea whether the client was still reading, and a dropped connection threw away everything that had already been produced. This note is about the plumbing that turns a demo stream into one you can leave running on a serverless bill without watching it. ##### Why the first token matters more than the last A model that takes eight seconds to finish a reply feels fast if the first word arrives in four hundred milliseconds and feels broken if the whole thing lands at once after eight seconds. Total latency is the same, but the reader is doing something during the wait in the first case and staring at a spinner in the second. Time to first token is the metric that tracks how the interface feels; time to last token only tracks cost. Streaming also changes what an error looks like. If the upstream fails after producing two paragraphs, a streamed client already has those two paragraphs on screen and can say so, while a buffered client has to show nothing and apologise. The trade is that a stream is a long-lived resource with two ends, and every problem in this note comes from one end not knowing what the other is doing. ##### Three transports, one decision There are three reasonable ways to move tokens to a browser. WebSockets give a bidirectional channel, which sounds ideal until you notice that a chat reply is entirely one directional: the client sends one request, the server sends many chunks, and the only message the client ever wants to send mid-stream is "stop". WebSockets also do not work on serverless route handlers without a separate long-lived process, and they bypass the HTTP caching, compression and proxy behaviour that everything else on the site relies on. Chunked fetch, meaning a plain `Response` whose body is a `ReadableStream`, is the simplest option and is what Zenith shipped with first. The client calls `fetch`, reads `response.body` with a reader, and appends text as it arrives. It has no framing, so the client cannot tell a heartbeat from content, and it has no event identity, so there is nothing to resume from. Server sent events sit in the middle. They are still a single HTTP response, so they run on any route handler and pass through any proxy that does not buffer, but the body has a line-based framing with `id`, `event` and `data` fields, comments for heartbeats, and a built-in `Last-Event-ID` header the browser sends on reconnect. The native `EventSource` API only supports GET, which is awkward for a request carrying a conversation, but nothing stops you from writing SSE frames from a POST handler and parsing them yourself with `fetch`. That is what the rest of this note does: SSE framing over a POST, parsed by hand, with the abort signal doing the work that a WebSocket's second direction would have done. ##### A route handler that streams and stops The handler below is the shape Zenith's endpoint converged on. Three details carry all the weight. The upstream call receives its own `AbortController`, and that controller is wired to `req.signal`, which the runtime aborts when the client disconnects. The stream uses `pull` rather than pushing everything from `start`, so the runtime only asks for the next upstream chunk when there is room in the queue. And `cancel` aborts upstream and returns the iterator, which releases the HTTP connection to OpenRouter instead of leaving it open until the model finishes on its own. ```ts import OpenAI from 'openai'; export const runtime = 'nodejs'; const MODEL = 'google/gemini-2.5-flash-lite'; const client = new OpenAI({ baseURL: 'https://openrouter.ai/api/v1', apiKey: process.env.OPENROUTER_API_KEY }); const encoder = new TextEncoder(); const frame = (id: number, data: string) => encoder.encode(`id: ${id}\ndata: ${JSON.stringify(data)}\n\n`); export async function POST(req: Request) { const { messages } = (await req.json()) as { messages: { role: 'user' | 'assistant'; content: string }[] }; const upstream = new AbortController(); req.signal.addEventListener('abort', () => upstream.abort(), { once: true }); const completion = await client.chat.completions.create( { model: MODEL, messages, stream: true, max_tokens: 500 }, { signal: upstream.signal }, ); const iterator = completion[Symbol.asyncIterator](); let seq = 0; let heartbeat: ReturnType | undefined; const body = new ReadableStream( { start(controller) { heartbeat = setInterval(() => { try { controller.enqueue(encoder.encode(': ping\n\n')); } catch { clearInterval(heartbeat); } }, 15_000); }, async pull(controller) { try { const { value, done } = await iterator.next(); if (done) { clearInterval(heartbeat); controller.enqueue(encoder.encode(`id: ${seq}\nevent: done\ndata: {}\n\n`)); controller.close(); return; } const text = value.choices[0]?.delta?.content ?? ''; if (text) controller.enqueue(frame(++seq, text)); } catch (err) { clearInterval(heartbeat); controller.error(err); } }, async cancel() { clearInterval(heartbeat); upstream.abort(); await iterator.return?.(); }, }, new CountQueuingStrategy({ highWaterMark: 8 }), ); return new Response(body, { headers: { 'Content-Type': 'text/event-stream; charset=utf-8', 'Cache-Control': 'no-cache, no-transform', 'X-Accel-Buffering': 'no', }, }); } ``` The `no-transform` directive and the `X-Accel-Buffering` header are there because the most common reason a stream "does not work in production" is a proxy or compression layer holding the response until it is complete. If the first token arrives on localhost and not on the deployed site, look there before touching the code. ```figure streaming-llm-responses/abort-propagation Tokens travel from the model through the handler to the browser, and a single abort signal travels the other way through the same three hops. ``` The abort path is worth tracing by hand. The visitor presses stop or navigates away. The client's `AbortController` fires, `fetch` rejects with an `AbortError`, and the browser closes the connection. On the server, the runtime observes the closed socket and calls `cancel` on the body stream and aborts `req.signal`. Either of those aborts the upstream controller, the OpenAI client's pending `fetch` rejects, and the model provider sees the connection close and stops billing tokens. Miss any one link and the chain breaks silently: the handler keeps pulling, the provider keeps generating, and nobody is reading. ##### Backpressure: when the reader is slower than the model A `ReadableStream` has an internal queue with a high water mark. When the queue holds fewer chunks than that mark, `desiredSize` is positive and the runtime calls `pull` to fetch more. When the queue is full, `desiredSize` drops to zero or below and the runtime stops calling `pull` until the consumer has read something. That is the entire backpressure mechanism, and it only works if the producer waits for `pull` instead of pushing. The version that pushes from `start`, looping `for await` over the upstream and enqueueing every chunk, ignores the mark completely. The queue grows to whatever the model produces, in memory, per request. For a five hundred token reply that is a few kilobytes and nobody notices. For a long generation being read by a phone on a bad connection, the handler holds the entire reply while the client drains it at its own pace, and the upstream connection stays open the whole time because the loop never blocks. ```figure streaming-llm-responses/backpressure When the queue reaches its high water mark, pull stops being called and the upstream iterator waits; the model is paused by the reader. ``` With `pull`, a slow reader pauses the model. The upstream HTTP response has its own flow control, so when the handler stops reading from the iterator, TCP receive windows close, the provider's send buffer fills and its generation loop blocks on write. The token cost is still incurred for whatever has been produced, but the handler's memory is bounded by eight chunks and the queue in the provider's socket, not by the length of the reply. The high water mark of eight is a judgement call: small enough that a slow reader is felt quickly, large enough that a burst of short deltas does not cause a `pull` round trip per token. ##### Heartbeats, reconnection and Last-Event-ID Idle connections get killed. Load balancers, mobile carriers and browsers all have timeouts, and a model that is thinking for twenty seconds before its first token looks idle to every one of them. The `: ping` comment line every fifteen seconds is invisible to the SSE parser but keeps bytes moving so the intermediaries see a live response. Fifteen seconds is below the common thirty second idle limits and above anything that would matter for bandwidth. When the connection does drop, the client wants to continue rather than start over, both for the reader and for the bill, because regenerating a reply that was ninety percent delivered is the most expensive way to finish it. SSE's answer is the `id` field on every frame. The client remembers the last `id` it processed and sends it back as `Last-Event-ID` on reconnect. The server then needs to be able to replay from that offset, which means the server needs to have kept the chunks somewhere the reconnecting request can reach them. ```figure streaming-llm-responses/resume-offset After the connection drops at event 7, the client reconnects with Last-Event-ID: 7 and the server replays from the stored buffer starting at 8. ``` On a single long running server that store is a `Map` from generation id to an array of chunks plus a done flag, and the handler tees each frame into it as it enqueues. On serverless it has to be external, because the reconnect may land on a different instance from the one still generating; a short lived key in Redis with a list per generation is enough. Either way the generation continues independently of the connection that started it, and a reconnecting client is served from the buffer until it catches up, then from live frames. ##### Partial output and idempotent resume Resume only works if the generation itself is idempotent from the client's point of view. Two rules make that true. The client never sends the same conversation twice as a new request while it holds a generation id; it sends the id and an offset instead. And the server never starts a second generation for an id it already knows; it either replays the buffer, or if the buffer is complete, sends the remaining frames followed by `done`. The generation id is minted by the client before the first request so a request that fails before the response headers arrive can be retried with the same id, and the server treats the second arrival as a resume from zero rather than a duplicate. This is the same idempotency key pattern used for payments and for tool calls in an agent loop, applied to an output that arrives in pieces. The offset is the `Last-Event-ID`, the key is the generation id, and the fingerprint is the conversation hash, so a client that reuses an id with different messages gets an error rather than someone else's reply. ##### Rendering incremental markdown without breaking the page The model writes markdown, and markdown is not designed to be parsed a token at a time. A stream that has delivered three backticks and half a line of TypeScript is inside an unclosed code fence; a naive renderer will flip the entire remainder of the message into a code block and flip it back a moment later. The fix is to render the whole accumulated string on each update, not the delta, and to give the renderer a copy with any open fence or open inline code closed with a synthetic terminator before parsing. The extra closing fence is discarded on the next update when the real one arrives. Two more details keep the client honest. Updates are batched with `requestAnimationFrame` so a burst of thirty tiny deltas causes one React render, not thirty. And the renderer is fed text, never HTML, with raw HTML disabled in the markdown parser, because the model's output is untrusted input that arrives on the same origin as the site; a model that can be convinced to emit a script tag must not be able to run it. ##### Cost control: stop generating when nobody is listening Everything above reduces to one rule: a token that no one will read should never be generated. The client hook below is the other end of the abort chain. It owns one `AbortController` per request, aborts the previous one when a new message is sent, aborts on unmount, parses SSE frames on the blank line delimiter, tracks the last event id for resume, and coalesces updates per animation frame. ```ts 'use client'; import { useCallback, useEffect, useRef, useState } from 'react'; type Status = 'idle' | 'streaming' | 'done' | 'error'; type Msg = { role: 'user' | 'assistant'; content: string }; type Frame = { id?: number; event?: string; data?: string }; function parseFrame(raw: string): Frame | null { const out: Frame = {}; for (const line of raw.split('\n')) { if (!line || line.startsWith(':')) continue; const sep = line.indexOf(':'); const field = sep < 0 ? line : line.slice(0, sep); const value = sep < 0 ? '' : line.slice(sep + 1).trimStart(); if (field === 'id') out.id = Number(value); else if (field === 'event') out.event = value; else if (field === 'data') out.data = (out.data ?? '') + JSON.parse(value); } return out.id === undefined && !out.event && out.data === undefined ? null : out; } export function useStreamedReply(url: string) { const [text, setText] = useState(''); const [status, setStatus] = useState('idle'); const controller = useRef(null); const lastEventId = useRef(0); const stop = useCallback(() => controller.current?.abort(), []); useEffect(() => stop, [stop]); const send = useCallback(async (messages: Msg[]) => { stop(); const ac = new AbortController(); controller.current = ac; lastEventId.current = 0; setText(''); setStatus('streaming'); let pending = ''; let raf = 0; const flush = () => { raf = 0; const chunk = pending; pending = ''; setText(t => t + chunk); }; try { const res = await fetch(url, { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: JSON.stringify({ messages }), signal: ac.signal, }); if (!res.ok || !res.body) throw new Error(`upstream ${res.status}`); const reader = res.body.pipeThrough(new TextDecoderStream()).getReader(); let buffer = ''; let finished = false; while (!finished) { const { value, done } = await reader.read(); if (done) break; buffer += value; let cut = buffer.indexOf('\n\n'); while (cut >= 0 && !finished) { const ev = parseFrame(buffer.slice(0, cut)); buffer = buffer.slice(cut + 2); cut = buffer.indexOf('\n\n'); if (!ev) continue; if (ev.id !== undefined) lastEventId.current = ev.id; if (ev.event === 'done') finished = true; else if (ev.data) { pending += ev.data; if (!raf) raf = requestAnimationFrame(flush); } } } if (raf) cancelAnimationFrame(raf); if (pending) flush(); setStatus('done'); } catch (err) { if (raf) cancelAnimationFrame(raf); setStatus((err as Error).name === 'AbortError' ? 'done' : 'error'); } }, [url, stop]); return { text, status, send, stop, lastEventId }; } ``` The `useEffect` that returns `stop` is the line most often missing from hooks like this. Without it, a visitor who opens the chat, asks a question and closes the panel leaves a fetch running that keeps the server pulling and the provider generating until the reply finishes. With it, unmount aborts the fetch, the socket closes, `cancel` runs on the server, and the upstream request ends within a round trip. On a widget that is opened out of curiosity far more often than it is read to the end, that one line is the difference between paying for the replies people read and paying for the replies they walked away from. ### How a TOTP Authenticator Really Works: HMAC, Time Windows and the Mistakes That Break Them - Tags: Security, Cryptography, TypeScript, Authentication - Summary: From the otpauth QR code to an accepted six digit code: HMAC-SHA1 over a counter, dynamic truncation, the thirty second step, verification windows, replay protection, constant time comparison and the base32 edge cases, with AdiNox as the worked example. #### How a TOTP Authenticator Really Works: HMAC, Time Windows and the Mistakes That Break Them AdiNox is an authenticator app Adil built in React and TypeScript: it scans the QR code a service shows during two-factor setup, stores the account locally, generates the six digit codes, and puts a face-api.js biometric check in front of the list. Building it forces you to understand the parts of TOTP that the standard makes look trivial and that real implementations get wrong in the same handful of ways. This note walks the whole path from the QR code to an accepted code, with the verifier side included, because an authenticator is only half of the system and the half that fails silently is usually the other one. ##### The secret is the whole system When a service enables two-factor authentication it generates a random secret, usually 10 to 32 bytes, and shows it to you once as a QR code. The QR code encodes an `otpauth://` URI: ```text otpauth://totp/Example:adil@example.com?secret=JBSWY3DPEHPK3PXP&issuer=Example&algorithm=SHA1&digits=6&period=30 ``` Everything after that is deterministic. Both sides hold the same bytes, both sides know the time, and a code is nothing more than a function of those two inputs. There is no server call when the app shows you a code and no network involvement when it changes; the service simply computes the same function when you submit and checks whether the results match. That is the elegant part. The consequence is that anyone holding the secret can compute every code that will ever be valid for that account, which is why the rest of this note keeps returning to where the secret lives and who can read it. The URI parameters are advisory in practice. The label and issuer are display strings. `algorithm`, `digits` and `period` are defined by the format but some widely used authenticator apps ignore them and assume SHA1, six digits and thirty seconds, so services that want to work everywhere use those defaults and services that deviate should test against the apps their users actually install. ##### HOTP: an HMAC over a counter TOTP is a thin layer on HOTP, and HOTP is a thin layer on HMAC-SHA1. Given the secret `K` and an 8 byte big-endian counter `C`, compute `HMAC-SHA1(K, C)` to get 20 bytes. The last nibble of the last byte is an offset between 0 and 15. Take the 4 bytes starting at that offset, clear the top bit so the result is a positive 31 bit integer, and reduce it modulo ten to the power of the digit count. The result is padded with leading zeros. That truncation step is called dynamic truncation and it is the part everyone reimplements from memory and gets subtly wrong. ```figure totp-internals-authenticator/hotp-pipeline The counter is signed with the secret, the last nibble of the tag selects four bytes, the top bit is cleared and the value is reduced modulo one million. ``` ```ts export async function hotp(secret: Uint8Array, counter: bigint, digits = 6): Promise { const key = await crypto.subtle.importKey('raw', secret, { name: 'HMAC', hash: 'SHA-1' }, false, ['sign']); const message = new Uint8Array(8); new DataView(message.buffer).setBigUint64(0, counter); const mac = new Uint8Array(await crypto.subtle.sign('HMAC', key, message)); const offset = mac[19] & 0x0f; const binary = ((mac[offset] & 0x7f) << 24) | (mac[offset + 1] << 16) | (mac[offset + 2] << 8) | mac[offset + 3]; return (binary % 10 ** digits).toString().padStart(digits, '0'); } ``` Three things in that function are load bearing. The counter is 8 bytes, big-endian, unsigned; an implementation that packs it as a 4 byte little-endian integer produces codes that agree with nothing. The mask `0x7f` on the first selected byte exists because the reference implementation was written in Java, where a 32 bit integer is signed and a set top bit would make the modulo negative; in JavaScript the bitwise operators are also 32 bit signed, so the mask is required here for exactly the same reason. And `padStart` matters because `000482` is a valid code, and roughly one in ten codes has a leading zero. A verifier that parses the submitted code as an integer will accept `482` for it, which is a sign that the comparison is not doing what the author thinks. SHA1 is broken as a collision resistant hash and that does not matter here: HMAC depends only on the compression function behaving like a pseudorandom function, and no practical attack on HMAC-SHA1 exists. SHA256 and SHA512 need only the hash name changed, but SHA1 is the default every app supports. ##### TOTP: the counter is the clock TOTP replaces the counter with the number of thirty second intervals since the Unix epoch. `C = floor(now / 30)`. That is the entire difference between the two algorithms. The device and the server each compute `C` from their own clock, sign it, truncate it and compare. At 12:00:00 the counter changes and the code changes with it; at 12:00:29 the same code is still valid; at 12:00:30 the next one is. The ring that counts down in an authenticator is thirty minus the remainder of the current second divided by thirty, and it is purely cosmetic: the code does not weaken as the ring shrinks. It does mean that a code copied at 12:00:29 and submitted at 12:00:31 has already expired on a strict server, which is why verifiers accept a window. ##### Windows, drift and what a window actually costs Two clocks never agree exactly. Phones sync with the network and are usually within a second or two, but a device that has been offline, a laptop with a dead clock battery, or a user typing slowly all produce codes for a step the server is no longer in. RFC 6238 suggests the verifier accept the current step and a small number of adjacent steps, typically one on each side. That turns a thirty second lifetime into a ninety second acceptance span. ```figure totp-internals-authenticator/time-window The server accepts the current step and one step either side, so a device whose clock runs slightly behind still produces a code inside the window. ``` The window is a security parameter and should be treated as one. Each extra step is another valid six digit code at any moment, so a window of one either side triples the number of codes an attacker guessing blindly will be accepted with, from one in a million to three in a million per attempt. That is still tiny, and the real defence against guessing is rate limiting on the verifier, not the window size; a service that allows unlimited attempts will fall to brute force with any window, and one that locks after a handful of failures is safe with a generous one. ##### Replay: the code you already accepted A code is valid for the whole step, plus the window. Nothing in the arithmetic stops someone who observed a code being submitted from submitting it again a few seconds later. Shoulder surfing, a phishing page that relays the code in real time, or a log line that captured the form body all produce a valid, unexpired code in an attacker's hands. The only defence is state: the verifier must remember what it has already accepted and refuse to accept it again. The cleanest form of that state is one number per account: the highest counter value for which a code has been accepted. On a successful verification, store the matched counter. On every verification, skip any step at or below the stored value before computing HMACs. This rejects the exact code that was just used, and it also rejects any older code in the window, which closes the case where a relayed code from the previous step is played after the legitimate one. ```figure totp-internals-authenticator/replay-cache The first submission of a code is accepted and its counter recorded; the same code submitted again inside the window is rejected before the HMAC is computed. ``` ```ts const STEP_SECONDS = 30; export function totpCounter(unixSeconds: number, step = STEP_SECONDS): bigint { return BigInt(Math.floor(unixSeconds / step)); } export async function totp(secret: Uint8Array, unixSeconds = Date.now() / 1000): Promise { return hotp(secret, totpCounter(unixSeconds)); } function equalConstantTime(a: string, b: string): boolean { if (a.length !== b.length) return false; let diff = 0; for (let i = 0; i < a.length; i++) diff |= a.charCodeAt(i) ^ b.charCodeAt(i); return diff === 0; } export type VerifyResult = { ok: true; skew: number } | { ok: false }; export class TotpVerifier { private lastAccepted = new Map(); constructor(private readonly window = 1) {} async verify(accountId: string, secret: Uint8Array, code: string, unixSeconds = Date.now() / 1000): Promise { if (!/^\d{6}$/.test(code)) return { ok: false }; const base = totpCounter(unixSeconds); const floor = this.lastAccepted.get(accountId) ?? -1n; let matched: { counter: bigint; skew: number } | null = null; for (let k = -this.window; k <= this.window; k++) { const counter = base + BigInt(k); if (counter <= floor) continue; const expected = await hotp(secret, counter); if (equalConstantTime(expected, code) && matched === null) matched = { counter, skew: k }; } if (matched === null) return { ok: false }; this.lastAccepted.set(accountId, matched.counter); return { ok: true, skew: matched.skew }; } } ``` The map lives in memory here; in a real service it is a column next to the secret, updated in the same transaction that marks the login as complete, so that two concurrent submissions of the same code cannot both pass a check that reads before either writes. A cache that expires entries after a minute or two is enough, because anything older than the window is rejected by the arithmetic anyway. ##### Constant time comparison `expected === code` short circuits on the first differing character. In principle a measurement of how long the comparison took leaks how many leading digits were right, and with six digits and a network in between that leak is far too small to exploit. The reason to use a constant time comparison anyway is that the habit costs nothing and the same function will eventually be used to compare something longer, like a session token or a webhook signature, where the leak is real. The loop above visits every character and folds the differences into one integer, so its running time depends only on the length, which is public. The verifier also computes every HMAC in the window before deciding, rather than returning on the first match. That keeps the number of HMAC evaluations independent of which step matched, which is the same discipline applied one level up. ##### Base32 is where implementations quietly break The secret in the URI is base32, RFC 4648 alphabet, `A` to `Z` then `2` to `7`. Base32 was chosen over hex or base64 because the alphabet has no lowercase letters, no `0`, `1`, `8` or `9` and no punctuation, which means a person can read it aloud or type it by hand from a printed backup without ambiguity. The things that break decoders are all at the edges of that alphabet. Services often display the secret in lowercase or in groups of four separated by spaces for readability, so the decoder must uppercase and strip whitespace. Padding with `=` is optional in URIs and mandatory in the RFC, so the decoder must accept both. And a secret of 10 bytes encodes to 16 characters exactly, but a 20 byte secret encodes to 32 characters and a 32 byte secret to 52 characters with 4 bits left over, so the decoder must discard trailing bits rather than treating them as a partial byte or an error. ```ts const ALPHABET = 'ABCDEFGHIJKLMNOPQRSTUVWXYZ234567'; export function base32Decode(encoded: string): Uint8Array { const clean = encoded.toUpperCase().replace(/[\s=]/g, ''); const out: number[] = []; let bits = 0; let value = 0; for (const ch of clean) { const index = ALPHABET.indexOf(ch); if (index < 0) throw new Error(`invalid base32 character: ${ch}`); value = (value << 5) | index; bits += 5; if (bits >= 8) { bits -= 8; out.push((value >>> bits) & 0xff); value &= ~(-1 << bits); } } return Uint8Array.from(out); } ``` The mask at the end of each byte keeps `value` to the bits not yet emitted so it never grows past 12 bits, which keeps the left shift well inside the 32 bit range that JavaScript's bitwise operators actually use. Drop that line and the decoder still works for short secrets and silently corrupts long ones, which is the kind of bug that passes a unit test with the sample key from the RFC and fails for one user in fifty. ##### Storing the secret on the device An authenticator holds a list of secrets that are each equivalent to a permanent second factor for an account. AdiNox keeps them in local persistence on the device and gates the interface behind a face-api.js check, so the code list is not visible to whoever picks up an unlocked phone. It is worth being precise about what such a gate protects against. A biometric check that runs in the app controls who can open the list through the app; it does not by itself change what is stored on disk. The threat it addresses is a person with physical access to the unlocked device, and it addresses that threat well. The threat it does not address is a process that reads the storage directly, and for that the secrets need to be encrypted at rest with a key the app cannot recover without the platform's own unlock, which on a phone means the OS keystore bound to the device credential. ##### TOTP is not a passkey TOTP proves that whoever is typing holds the secret, and nothing more. It does not know which site is asking. A phishing page that looks like the real login can collect the password and the current code, relay both to the real service within the thirty second window, and log in; replay protection stops the second use of that code but not the first, because the attacker's use is the first. TOTP is also a shared secret, so a breach of the service's database leaks every user's second factor at once. Passkeys, meaning WebAuthn credentials, fix both. The private key never leaves the device, the server holds only a public key, and the signature the device produces covers the origin of the page that requested it, so a credential registered at the real domain is useless to a lookalike. That is why login is moving to passkeys. TOTP remains what it has always been: a cheap, offline, standardised second factor that works on any device with a clock, that a person can carry on paper if they must, and that is far better than a password alone provided the verifier gets the window, the replay cache and the comparison right. ### Recon for Juicy Endpoints via URLScan Dorking - Tags: URLScan, Hacking, Pentesting, Bug Bounty - Summary: A methodology for authorized testers: narrowing a huge in scope estate down to the few hosts worth a human's time through passive discovery, polite confirmation, and triage, plus the defensive mirror for asset owners. #### Recon for legacy endpoints: focusing a large scope in authorized testing > **Scope and ethics.** This is a methodology write-up for authorized penetration testers and bug-bounty participants. Everything here assumes an explicit, written engagement that lists the assets and the rules. I do not test anything I am not permitted to, and neither should you. The value is in the method (how to spend limited time on the right surface), not in any single payload. Wildcard programs and enterprise engagements hand you a huge in scope surface: thousands of subdomains across a main brand and its acquisitions. Blindly crawling all of it wastes the one resource you cannot get back: time. The skill is triage: narrowing thousands of hosts down to the handful worth a human's attention, all within scope. ##### The insight that drives the triage Modern single-page apps ship with sensible defaults, so the interesting weaknesses tend to cluster around older technology: legacy extensions (`.php`, `.aspx`, `.jsp`) served by dated stacks, forgotten admin panels, and debug endpoints left enabled. These are not guaranteed to be vulnerable. They are simply where the base rate is higher, so they earn the first look. ```mermaid flowchart LR A[In-scope wildcard] --> B[Passive discovery] B --> C[Filter to legacy tech] C --> D[Confirm live hosts] D --> E[Manual triage and prioritize] E --> F[Authorized testing within scope] ``` ##### Phase 1: passive discovery first Start with sources that do not touch the target. Passive datasets (search-oriented scanners like URLScan.io, certificate transparency logs, and public subdomain data) let you understand the estate before sending a single packet. Filtering these for legacy extensions and telltale titles (login pages, admin panels) turns an unmanageable list into a shortlist. Passive-first is both faster and lower-impact, which matters when real production systems are in scope. The output of this phase is not "targets to attack" but "hosts worth confirming", typically a few dozen out of thousands. ##### Phase 2: confirm what is actually live Public data goes stale. Before spending attention on a host, confirm it still resolves and responds, and note its technology stack. Lightweight, well-behaved probing here (respecting rate limits and the program's rules) separates the live, relevant hosts from historical noise. Tooling from the ProjectDiscovery ecosystem (subfinder for discovery, httpx for liveness, katana for scoped crawling) is common for this, configured conservatively so you are a good citizen on someone else's infrastructure. The principle: be quiet and slow. Aggressive scanning is both noisy and, on production systems, potentially disruptive, neither of which serves the engagement. ##### Phase 3: triage by likely attack surface This is where judgment beats automation. With a confirmed, in scope shortlist, score each host by what it plausibly exposes, and let that ordering decide where a human spends time. | Endpoint pattern | Priority | Class of concern to review | | --- | --- | --- | | Admin or login panels | Critical | Authentication and access control | | File upload or processing endpoints | High | Input handling and file validation | | Parameterized search or lookup views | Medium | Injection classes, output encoding | | Error and debug pages | Low | Information disclosure, stack traces | The table intentionally names *classes of concern*, not exploits. The tester's job is to verify, within scope, whether a control is present or missing (for example, whether a login endpoint enforces rate limiting and proper session handling), and to document it clearly. ##### Phase 4: verify carefully, document precisely Automated scanners are good at breadth and bad at judgment, so use them to surface candidates and then confirm by hand. Nuclei and similar template-driven tools can flag known issues across the shortlist; a human then validates each finding, discards false positives, and writes it up. A finding is only useful to the asset owner if it comes with clear reproduction steps, an accurate severity, and a concrete remediation. Throughout, stay inside the lines: only in scope hosts, only the tests the program permits, and no action that could degrade a production service. If a finding would require crossing scope to confirm, the correct move is to report the exposure and stop, not to push further. ##### What good recon actually optimizes for It is tempting to measure recon by volume (a million URLs collected), but the metric that matters is *signal per hour*. A tester who narrows ten thousand hosts to five well-chosen ones, verifies them carefully, and writes them up cleanly delivers more value than one who floods the program with low-quality, automated noise. Programs and their triagers reward precision. ##### Defensive mirror: what this means for asset owners If this pipeline finds your systems easily, so will less friendly parties. The same method points straight at the defenses worth having: - **Know your estate.** Maintain an inventory of subdomains and retire the ones you no longer run. Forgotten hosts are the ones that get found. - **Decommission legacy stacks.** Old frameworks on unpatched servers are the highest-base-rate targets; migrate or remove them. - **Disable debug in production.** Verbose errors and debug endpoints hand attackers version and path information for free. - **Monitor for enumeration.** Passive discovery is invisible, but active probing is not. Watch for it, and rate-limit accordingly. ##### Takeaways Effective reconnaissance is triage, not collection. Discover passively, confirm politely, prioritize by likely surface, verify by hand, and stay rigorously within scope. Done this way it is a professional, low-impact discipline that helps owners find their weak spots before someone with worse intentions does, which is the entire point of authorized testing. ### How Attackers Exploit Public WiFi to Sniff Decrypted Network Packets - Tags: WiFi Security, Hacking, Pentesting - Summary: How public-WiFi attacks actually work (evil-twin access points, on-path interception, TLS downgrade) and why end-to-end encryption, HSTS preload, encrypted DNS, and a VPN make an attacker's position on the wire close to worthless. #### How attackers exploit public WiFi, and how to defend against it > **Scope and ethics.** This is a defensive explainer for engineers, network owners, and authorized testers. It describes attacker methodology at the level needed to defend against it, not a set of operational instructions. Only assess networks you own or are explicitly authorized to test, and stay within the agreed scope and local law. Open and lightly protected WiFi (cafes, airports, hotels) is attractive to attackers because it removes the usual barrier of getting onto the same network as the victim. Understanding how that plays out lets you design networks and applications that stay safe even when the link layer cannot be trusted. The short version: treat every network you did not build as hostile, and push security up the stack. ##### Why public WiFi is a soft target A few properties combine to make these networks risky: - **Weak or absent link encryption.** Open networks send frames in the clear; shared-password networks give every client the same key, so being on the network is enough to be in a position to observe others. - **A shared medium.** WiFi is broadcast by nature. Devices on the same access point are, by default, near each other's traffic. - **Trusting clients.** Devices happily reconnect to any access point advertising a familiar network name, and users click through certificate warnings. None of these are exotic bugs; they are the default behavior of the technology, which is exactly why the defense has to live above it. ##### The attacker's playbook, described defensively The point of naming these techniques is to know what each one relies on, so you can remove that reliance. | Technique | What it relies on | What removes the leverage | | --- | --- | --- | | Rogue / "evil twin" access point | Clients auto-trusting a network name | VPN and TLS make the captured traffic useless; user education | | On-path interception (ARP spoofing) | Plaintext or downgradeable traffic | HTTPS everywhere, HSTS preload, encrypted DNS | | Passive sniffing | Unencrypted application protocols | Encrypt every protocol; retire plain HTTP, FTP, Telnet | | TLS downgrade / stripping | Redirect-based first hops over HTTP | HSTS preload, so the first request is already HTTPS | | Handshake / key attacks on the WiFi layer | Weak pre-shared keys, unpatched clients | WPA3 or enterprise auth, patched devices, strong keys | The pattern is consistent: an attacker who controls the network can see and sometimes alter traffic, but strong end-to-end encryption turns everything they capture into noise. ##### The defense, from strongest to supporting **1. End-to-end encryption is the real boundary.** If every connection is TLS with a valid certificate, an attacker on the path sees ciphertext and metadata, not content. For applications, that means HTTPS with no plaintext fallback and certificates validated strictly: no "ignore this warning" affordances. **2. HSTS preload closes the downgrade gap.** Stripping attacks depend on that first, pre-redirect request going out over HTTP. Being on the HSTS preload list means the browser refuses HTTP for your domain from the very first request, before any network attacker can intervene. ``` # Send this on every HTTPS response, then submit the domain to the preload list. Strict-Transport-Security: max-age=63072000; includeSubDomains; preload ``` **3. Encrypted DNS.** DNS over HTTPS or over TLS stops an on-path observer from reading or tampering with name resolution, removing an easy redirection and surveillance channel. **4. A VPN on untrusted networks.** A VPN tunnels all traffic to a trusted egress, so a hostile access point sees only the encrypted tunnel. It is the most reliable single control for a user who has to work from an unknown network. **5. Fix the link layer where you own it.** If you run the network, prefer WPA3 or WPA2-Enterprise with per-user credentials rather than a single shared password, use strong keys, isolate clients from one another, and keep access-point firmware current. **6. Pin certificates in native apps.** Web browsers rely on the public certificate authority system, which an attacker with a trusted rogue certificate can sometimes abuse. Mobile and desktop apps do not have to: pinning the expected certificate or public key means the app rejects any TLS session that does not present the exact identity you shipped, closing the door on interception even when the device has been coaxed into trusting an attacker's certificate. The cost is operational (you have to rotate pins carefully before certificates expire), but for high-value apps it is worth it. ##### For the authorized tester A legitimate wireless assessment is about demonstrating and measuring exposure, not maximizing harm. Good practice looks like: - A written, scoped authorization that names the networks and time window. - An isolated test environment or a clearly bounded engagement, so real users are never swept up. - Reporting that translates findings into fixes: which protocols were observable, which clients trusted a spoofed network, which services lacked HSTS. - Recommendations the owner can act on (enterprise authentication, HSTS preload, VPN policy, client patching) rather than raw capture files with no context. ##### What to tell users Most real-world protection comes down to a few habits: prefer networks you trust, keep a VPN on when you cannot, never dismiss a certificate warning, and keep devices patched. Those steps assume the network is hostile, which on public WiFi is the only safe assumption. ##### Takeaways Public WiFi attacks work by controlling a network the victim did not build. You cannot always control the network, so you make it not matter: encrypt end to end, preload HSTS so the first request is already protected, encrypt DNS, and tunnel over a VPN on anything untrusted. Defenders who move security up the stack make the attacker's position on the wire close to worthless, which is the whole goal. ### Securing In-Game Economies Against Value Tampering - Tags: Hacking, Game Security, Reverse Engineering - Summary: Where game economies get tampered with (memory edits, forged packets, race conditions, injection) and how server-side authority, atomic transactions, and audit logging shut it down. A defensive guide. #### How in-game economies get tampered with, and how to defend them > **Scope and ethics.** This is a defensive write-up for game developers and authorized security testers. Everything below is about understanding attacker methodology so you can design against it. Only ever test systems you own or have explicit, written permission to assess, and keep to the agreed scope. Online games run an economy (coins, currency, items), and that economy is only as trustworthy as the boundary between the player's machine and your server. When teams get value tampering wrong, it is almost always because the client was trusted with something the server should have owned. This post walks through where that trust breaks and how to close the gaps, without providing a recipe for cheating. ##### The one rule that prevents most of this **The client is hostile. The server is authoritative.** Anything the player's device reports (score, balance, position, "I bought this") is a request, not a fact. If your server writes a currency value because the client told it to, you have already lost. Every economy-relevant change must be computed and validated server-side against rules the client cannot influence. Most published cheating techniques are just different ways of exploiting a server that forgot this rule. ##### Where attackers look, and what it tells defenders At a high level, tampering falls into a few categories. For each, the useful question is not "how is it done" but "what did the server fail to validate." | Attacker angle | Underlying weakness | Defensive answer | | --- | --- | --- | | Editing client memory or local state | Client treated as source of truth | Recompute balances server-side; never trust submitted totals | | Replaying or forging network messages | No integrity or freshness on messages | Authenticated sessions, nonces, sequence numbers | | Race conditions ("spend then re-add") | Non-atomic economy transactions | Serialize economy changes in a single transaction | | Injection against game or web endpoints | Unvalidated input reaching a query | Parameterized queries, strict input validation | | Tampered or repackaged clients | Reliance on client-side checks | Server-side authority plus integrity attestation | ##### Reconnaissance, from the defender's chair Authorized testers begin by mapping the attack surface the same way an attacker would, so it helps to know what they see. They enumerate the transport (which ports, TCP or UDP), watch traffic to understand the protocol, and inspect the client to learn what data it sends and how it is protected. The defensive takeaway from each step: - **Traffic is observable.** Assume anyone can read your protocol. Encrypt transport (TLS/DTLS), and never send authority-bearing values in the clear. - **The client is inspectable.** Anything shipped to the player, including "secret" checks, can be read and modified. Treat client-side validation as UX, never as security. - **Endpoints are discoverable.** Your economy APIs will be found. Authenticate and authorize every one of them. ##### Server-side authority in practice The defense is not exotic; it is discipline. A currency change should look like this on the server: ```python # Server computes the outcome; the client's numbers are inputs to validate, never the result. def purchase_item(session, item_id: str): user = session.authenticated_user # from a verified session, not the request body item = catalog.get(item_id) if item is None: raise InvalidRequest("unknown item") with db.transaction(): # atomic: no spend-then-readd races balance = wallet.balance_for_update(user.id) # row lock if balance < item.price: raise InsufficientFunds() wallet.debit(user.id, item.price) inventory.grant(user.id, item.id) audit.log(user.id, "purchase", item_id, item.price) ``` Three properties matter here: the identity comes from a verified session rather than the payload, the whole change is one atomic transaction so concurrent requests cannot interleave into free currency, and every change is written to an audit log. ##### Detection: assume some tampering will get through No boundary is perfect, so instrument the economy as if it will be probed: - **Rate and anomaly checks.** Flag currency gains that outrun any legitimate play pattern. A balance that jumps beyond what the game's rules can produce is a signal, not a mystery. - **Server-side invariants.** Periodically reconcile balances against the transaction log. If they disagree, something wrote currency outside the sanctioned path. - **Audit trails.** Every economy mutation should be reconstructable from logs. This is what turns "a player has too much gold" into "here is exactly which endpoint produced it." - **Integrity signals.** Client attestation and anti-tamper tooling raise the cost of a modified client, but treat them as friction that buys time, not as a wall. ##### Designing the economy to be boring to attack The most robust economies are the ones where cheating simply has nothing to grab: 1. Keep every authoritative value on the server. The client renders a balance; it never owns one. 2. Make each economy change atomic and idempotent, so replays and races produce no gain. 3. Validate and parameterize all input, on the assumption that any endpoint can be hit directly. 4. Encrypt transport and bind sessions to identity, so messages cannot be trivially forged or replayed. 5. Log and reconcile, so tampering is detectable after the fact even when prevention slips. ##### Takeaways Value tampering is a symptom, and the disease is misplaced trust. When the server computes outcomes, guards them in atomic transactions, and reconciles against an audit log, the client-side techniques attackers reach for lose their leverage: there is no local number worth editing, because the server was never going to believe it. Build the boundary correctly and cheating turns from a payday into a nuisance. ## Contact Available for new engagements; remote worldwide. Email is best; replies within a day. - Email: adilmunawarx@gmail.com (mailto:adilmunawarx@gmail.com) - WhatsApp: +92 324 4965220 — https://wa.me/923244965220 - LinkedIn: linkedin.com/in/adilmunawar — https://linkedin.com/in/adilmunawar - GitHub: github.com/adilmunawar — https://github.com/adilmunawar - Telegram: @adilmunawar — https://t.me/adilmunawar - Instagram: @adilmunawarx — https://instagram.com/adilmunawarx ## How to cite / contact - Cite as: Adil Munawar, "Machine-learning engineer for agricultural remote sensing · full-stack developer", https://adilmunawar.vercel.app (accessed via llms-full.txt, last updated 2026-09-29). - Canonical source for all facts above: https://adilmunawar.vercel.app. Concise summary: https://adilmunawar.vercel.app/llms.txt. Structured data: https://adilmunawar.vercel.app/ai-profile.json. - Do not attribute metrics, prices, dates or client names that are not stated here; the private model entries deliberately carry none. - For hiring or collaboration, email adilmunawarx@gmail.com or use the channels in the Contact section.