[{"content":"","date":"June 22, 2026","permalink":"/roasts/011-guatemala/","summary":"","title":"011 — Guatemala Huehuetenango Rolando Sanchez","type":"roasts"},{"content":"","date":"June 22, 2026","permalink":"/roasts/012-guatemala/","summary":"","title":"012 — Guatemala Huehuetenango Rolando Sanchez","type":"roasts"},{"content":"","date":"June 21, 2026","permalink":"/roasts/010-guatemala/","summary":"","title":"010 — Guatemala Huehuetenango Rolando Sanchez (Full City+, pushed longer)","type":"roasts"},{"content":"","date":"June 20, 2026","permalink":"/roasts/008-guatemala/","summary":"","title":"008 — Guatemala Huehuetenango Rolando Sanchez (Full City+)","type":"roasts"},{"content":"","date":"June 20, 2026","permalink":"/roasts/009-guatemala/","summary":"","title":"009 — Guatemala Huehuetenango Rolando Sanchez (Full City+)","type":"roasts"},{"content":"","date":"June 19, 2026","permalink":"/roasts/006-nicaragua/","summary":"","title":"006 — Nicaragua Buenos Aires Maracaturra","type":"roasts"},{"content":"","date":"June 19, 2026","permalink":"/roasts/007-nicaragua/","summary":"","title":"007 — Nicaragua Buenos Aires Maracaturra (Full City)","type":"roasts"},{"content":"","date":"June 14, 2026","permalink":"/roasts/005-guatemala/","summary":"","title":"005 — Guatemala Huehuetenango Unión Cantinil","type":"roasts"},{"content":"Sourced from the official 2000AB 120V Operation Manual V30 (Oct 2024) and the Behmor knowledge base . Pulled into the site so it\u0026rsquo;s available on-hand during a roast.\nThe single most important thing: P1–P5 mean TWO things This is the easiest way to wreck a roast on this machine: P1–P5 have different meanings depending on whether you press them before or after Start.\nWhen What P1–P5 means BEFORE Start (Auto Mode) Origin-specific roast profile. Choose based on bean type. AFTER Start (Manual Mode) Power level (0/25/50/75/100%). Use to drop power post-FC. The manual makes this explicit (Part 3, page 8 and Part 5 page 12). If you treat P5 like \u0026ldquo;100% power\u0026rdquo; when selecting it as your starting Auto profile, you\u0026rsquo;re actually picking the Hawaiian/island lowest-heat profile — wrong direction.\nAuto Mode — pick by origin BEFORE pressing Start This is the chart from the manual that drives every roast plan on the site.\nProfile Bean type Recommended origins Notes P1 Hard bean (SHB / HG, high-grown) All Centrals, Peruvian, Colombian Default profile. The manual recommends new users do four ¼ lb batches on P1 to learn the machine. P2 Hard bean (SHB / HG, high-grown) All Centrals, Peruvian, Colombian Less aggressive than P1 for the same regions; useful for a softer entry into FC on dense lots. P3 Soft bean / low-grown Brazilians, Africans, Southeast Asian, Malabar, Jamaican Blue Mountain, Yauco Selecto Add time via \u0026#43; or C button — P3–P5 auto programs may need extension. P4 Soft / low-grown island Kona, low-grown island coffees Used for taking Kona-style beans just into second crack. P5 Soft / low-grown island Kona, low-grown island coffees Lowest heat profile. Used for keeping island coffees shy of second crack at City/City\u0026#43;. Practical reading: Centrals / Peruvian / Colombian → P1 (or P2 for slightly gentler). Brazilian / African / SE Asian → P3. Kona and island coffees → P4 or P5. The default for the machine is P1; the manual recommends beginners run four ¼ lb batches on P1 to learn it.\nManual Mode — press P1–P5 AFTER Start Once a roast has started, the P buttons become power-level adjusters.\nButton Power Use OK after FC? P1 0% Stalls heat; lets exothermic momentum carry. Press 1 button to return to Auto. ✓ P2 25% Gentle finishing power. ✓ P3 50% Manual recommends pressing P3 ~10–15 sec after FC starts to drop power and let beans\u0026#39; exothermic momentum carry. Combine with D button. ✓ P4 75% Pre-FC drying / Maillard only. ⚠ Not recommended after first crack (manual: combined with bean-mass heat may scorch). ✗ P5 100% Pre-FC drying phase only. ⚠ Not recommended after first crack (manual: combined with bean-mass heat may scorch). ✗ P4 and P5 are explicitly flagged as not recommended after first crack — combined with the bean mass already heated, they will scorch.\nTwo post-FC techniques — pick by goal The manual documents one approach (Professional Sample Roasting); forum and Sweet Maria\u0026rsquo;s-library consensus uses a different one (C-button + Auto P1 hold) for drinking coffee. Both start the same way: ½ lb + P1 + Start. They diverge at first crack.\nProfessional Sample Roasting (manual Part 5, page 12-13) — sample roasting only Press desired weight + P1 + START (defaults to Auto P1) Once you hear first cracking, wait 10–15 seconds Press P3 (drops to 50% power in Manual mode) Press D (changes drum speed) May need to extend time using + or C button Press Cool when target hit The manual\u0026rsquo;s rationale: \u0026ldquo;Cutting power using the P3 button to 50% sustains heat, but gently allows the beans\u0026rsquo; exothermic momentum to work and prevent scorching beans.\u0026rdquo;\nUse this for: small evaluation batches where slow, low-development finish is intentional (cupping samples).\nDon\u0026rsquo;t use this for drinking coffee at ½ lb or 1 lb. Roasts 006 and 007 logged on this site both used this technique and came in under-developed (8–9.4% loss vs 13–16% target, DTR 38%+) — the P3 cut starves development heat over the 3–5 minute development window.\nC-button + Auto P1 hold — drinking coffee at ½ lb / 1 lb Synthesized from the Behmor knowledge base, Sweet Maria\u0026rsquo;s library, and home-barista forum consensus.\nPress desired weight + P1 + START At the 75% shutoff prompt: press C to continue (also bypasses Err 7) At first crack: check the remaining clock first. If there\u0026rsquo;s less than ~2:30 remaining, press C to reset to the Rosetta Stone duration. If there\u0026rsquo;s more, don\u0026rsquo;t press C — let Auto P1\u0026rsquo;s original program run. Stay on Auto P1 through the FC body — full heat carries development If you need more time near the end of the clock (low voltage, humid day): press C to extend toward the Rosetta Stone duration Only drop to P3 if 2C threatens (rapid louder snaps, acrid smoke, oil sheen forming) — as a brake, not as the default Press Cool when target hit Rationale: full Auto P1 heat through FC body actually drives off moisture and progresses Maillard. The C-button shapes the time without cutting power.\nUse this for: drinking coffee at ½ lb or 1 lb, targeting City through Full City+.\nPure Manual mode + 50/30/20 phase targets — community technique (not from the manual) Source: Virtual Coffee Lab YouTube — Behmor Coffee Roasting Recipe (li_XZd1hrtI , Sep 2021). Demonstrated on a Behmor 1600 Plus, not the 2000 AB Plus, so the heat-management button behavior and safety-shutoff threshold differ from this machine. Logged here as a community-sourced alternative to the two techniques above. See /behmor-video-notes.md at the repo root for the full breakdown.\nPreheat the chamber to ~210 °F (B-temp), then load drum Press desired weight + P1 + START, then immediately press P5 to switch to Manual mode (P-buttons now act as power levels, not Auto profiles) Log time + temperature every 30 seconds. Mark events: dry end, FC, drop Ride the P-buttons by chamber temperature — back off to P4 (75%) when approaching the safety ceiling; return to P5 if temp falls; press D for high drum + convection somewhere in the dry phase At first crack: drop to P3 (50%) for development Cool at the target level, accounting for ~15 s of cool-cycle carryover Target phase split (his framework, not the manual\u0026rsquo;s): 50% dry / 30% browning / 20% development of total elapsed time, with naturals running shorter on development. Hitting 53/32/14 in the video produced his preferred cup on the bean.\nCaveats before using this on a 2000 AB Plus:\nThe \u0026ldquo;back off to P4 to avoid the 320 °F shutoff\u0026rdquo; beat is a 1600+ behavior. The 2000 AB Plus uses the 75% Err 7 shutoff documented above — the trigger condition is different, so don\u0026rsquo;t carry the 4:15 / 295 °F intervention point across blindly The voiceover never mentions the C button, but Mike does press C near the end of the roast (observed in the video, unnarrated) — consistent with using C as a late extension lever per the + vs C table below. The C-button + Auto P1 hold technique above is still the project\u0026rsquo;s default for drinking coffee on this machine — try this one as a comparison, not as a replacement Drop is before 2C — this is a Medium / Full City recipe, not a Medium-dark target. For Full City+ on this machine, the C-button + Auto P1 hold path is better characterized + vs C — the most important machine quirk This is what trips up new Behmor users. + adds to the clock. C resets the clock. They are not interchangeable.\nButton What it does When to press + ADDS ~10s per press to the current clock Fine-tuning a small definite extension. Has a per-roast cumulative cap (~+2:00 observed). C RESETS the clock to the Rosetta Stone 1C→2C duration (does NOT add) When the current clock has LESS time than you need for development. If current clock has MORE, C shortens the roast. The C-button reset values (approximate) Weight setting Rosetta Stone reset value ¼ lb ~1:30 ½ lb ~2:09 (observed on roast 008) 1 lb ~3:00+ When pressing C HELPS vs HURTS Situation Press C? Why 75% shutoff (Err 7) fires Yes Bypasses Err 7. Reset value is roughly neutral here. FC arrives with under ~2:30 left on clock Yes C extends the clock to the Rosetta Stone value. FC arrives with 3:00+ left on clock (e.g. early FC) No C will SHORTEN the clock. Let Auto P1\u0026rsquo;s original program run. Near end of clock (under ~1:30) and you need more time Yes C extends. Use repeatedly in low-voltage situations. Failure mode logged on roast 008: FC came at 7:30 with 4:30 remaining on the ½ lb / 12:00 program. Pressing C at FC reset the clock to 2:09 — a 2:21 cut to available development time. Result: 8.85% loss vs 16–18% Medium-dark target, because the bean didn\u0026rsquo;t get enough post-FC time at full Auto P1 heat.\nThe forum technique \u0026ldquo;press C at FC\u0026rdquo; assumes FC arrives near the end of the program (typical for ½ lb dense SHB ~10–11 min FC). It doesn\u0026rsquo;t generalize to early-FC situations.\nWeight settings — default program time only The weight button sets the program duration. Heat is controlled by the profile (Auto) or by the P-power level (Manual).\nActual charge Weight button Program time 75% shutoff at Why ¼ lb (≈113 g) ¼ lb 8:30 6:22 elapsed Manual\u0026#39;s official beginner starting point: \u0026#34;NEVER START with quantities above ¼ lb\u0026#34; until familiar with the roaster. Start with four ¼ lb batches on P1. ½ lb (≈227 g) ½ lb 12:00 9:00 elapsed Move here once comfortable with ¼ lb roasts. 75% shutoff at 9 min elapsed — often coincides with FC. 1 lb (≈454 g) 1 lb 18:00 13:30 elapsed Full batch capacity. For Full City\u0026#43; darker, manual recommends REDUCING to 336 g (12 oz) for control — full pound retains too much thermal momentum past City\u0026#43;. Match the weight button to your actual charge. The 75% Err 7 safety shutoff fires at 75% of the program time. For Full City+ darker roasts, manual recommends reducing the batch to 336 g (12 oz) for control — full-pound batches retain too much thermal momentum past City+.\nPreheat sequence Weigh green into drum. Keep drum out of machine. Press any weight + any profile + START. Run for 1:30 minutes with drum out. Press OFF. Insert loaded drum + chaff tray. Start the actual roast. ⚠ Do not exceed 1:30 preheat — the manual warns that longer preheats trigger a safety feature that prevents starting the roast.\n75% safety shutoff (Err 7) At 75% through the program, the display blinks un:30..29..28 with loud beeping. You have 30 seconds to press START to continue. If you miss it, the machine enters cool mode and shows Err 7.\nThe C button bypasses Err 7 — pressing it during the warning window resets the timer per the Rosetta Stone and tells the machine you\u0026rsquo;re actively engaged.\nThe C button (Rosetta Stone) The C button RESETS the clock to an approximate 1C→2C duration for the selected weight setting (~1:30 for ¼ lb, ~2:09 for ½ lb, ~3:00+ for 1 lb). It does NOT add to the current clock.\nThis distinction matters a lot — see the \u0026ldquo;+ vs C\u0026rdquo; section below for when pressing C extends vs shortens a roast. The short version: C extends only when the current remaining clock is LESS than the Rosetta Stone value. If MORE, C shortens.\nThree legitimate uses:\nBypass the Err 7 safety prompt at 75% shutoff — signals active interaction Extend time near end of clock — Behmor KB notes \u0026ldquo;press it numerous times towards the end of the roast\u0026rdquo; for low-voltage extension Anticipate 2C — when used near end of program, the reset gives you a fresh count toward 2C A/B temperature buttons A — exhaust channel temperature (active once exhaust fan starts mid-roast) B — chamber wall temperature Manual is emphatic: \u0026ldquo;These temps ARE NOT bean temps and for information / correlation purposes only. Never gauge degree of roast based on these temps.\u0026rdquo;\nWhat \u0026ldquo;dark\u0026rdquo; means on this machine The level reference table in Methods \u0026amp; reference defines five levels from Light through Dark. Not all of them are safely targetable on the Behmor 2000 AB Plus. The manual is explicit on two limits:\n\u0026ldquo;The Behmor Roaster is not intended to roast coffee to levels known as Vienna, French, or Italian.\u0026rdquo; \u0026ldquo;Never roast past 10 seconds into second crack.\u0026rdquo; Mapping our level table to what\u0026rsquo;s achievable on this machine:\nLevel Drop trigger (per data/roast_guidance.yaml) Behmor-safe? Light 15–45s after FC starts, before any oil sheen Yes Light-medium 30–60s into development Yes Medium End of FC through 1:30–2:30 development Yes Medium-dark First few snaps of 2C Yes — at the safety ceiling. Drop within 10 sec of first 2C snap. Dark Rolling SC, before oil floods No — requires staying in 2C past the 10-second line; violates the manual. So in practice the deepest roast you can target safely is Medium-dark / Full City+, with the drop trigger being the first 1–3 snaps of 2C. For darker than that you need a different machine.\nCritical safety constraints These come straight from the manual and they are not soft rules:\nAllow 1 HOUR between roasts. Not minutes — the manual states this as a direct rule (Part 4 Important Reminders #4). Never roast past 10 seconds into second crack. Past this point beans can ignite. Never start above ¼ lb when first learning. The manual is loud about this — start with four ¼ lb batches on P1. The 2000 AB Plus is not intended for Vienna, French, or Italian roasts. Oily beans + chaff tray = fire risk. Never leave the roaster unattended. Explicit in the manual title: \u0026ldquo;NOT a set and walk away device.\u0026rdquo; Beginner guidance From manual Part 3:\nRoast four ¼ lb batches on P1 to learn FC sounds, color cues, smell changes. After ¼ lb is comfortable, move to ½ lb. Move to 1 lb last. Always forward-think the roast by 15 seconds (the reversal time from roasting to cooling). Never gauge degree of roast by oil appearance — drum roasters don\u0026rsquo;t show oil during the roast; droplets appear 24h–several days after, especially with rest. Cleaning Every 5 roasts (more often for darker roasts), do a dry burn (¼ lb + Start, run full program) with the interior cleaned by Simple Green or other non-caustic / non-abrasive cleaner. Never spray heating elements. Keep the back-lower-right corner of the roasting chamber clean — the thermocouple sensor lives there. Dirty sensor = extended / wrong roast times. Sources Behmor Roaster Quick Start Guide — Behmor knowledge base Control Panel Features \u0026amp; Auto-Safety Shutoff — Behmor knowledge base 2000AB Plus 120V Operation Manual V30 (Oct 2024) — official manual ","date":"June 14, 2026","permalink":"/roasts/behmor-2000ab/","summary":"Operational reference for the Behmor 2000 AB Plus pulled from the official Operation Manual V30 (Oct 2024) — origin-specific Auto Mode profiles, Manual Mode power controls, the pro sample roasting technique, safety, preheat, and full machine guidance.","title":"Behmor 2000 AB Plus — reference","type":"roasts"},{"content":"","date":"June 13, 2026","permalink":"/roasts/_template/","summary":"","title":"","type":"roasts"},{"content":"","date":"June 13, 2026","permalink":"/roasts/001-brazil/","summary":"","title":"001 — Brazil Machado Recanto do Engenho","type":"roasts"},{"content":"","date":"June 13, 2026","permalink":"/roasts/002-colombia/","summary":"","title":"002 — Colombia Sierra Nevada Finca La Esperanza","type":"roasts"},{"content":"","date":"June 13, 2026","permalink":"/roasts/003-nicaragua/","summary":"","title":"003 — Nicaragua Buenos Aires Maracaturra","type":"roasts"},{"content":"","date":"June 13, 2026","permalink":"/roasts/004-ethiopia/","summary":"","title":"004 — Ethiopia Kayon Mountain Taaroo","type":"roasts"},{"content":"Every roast page computes derived values from a few raw inputs and compares them against industry-consensus targets. This page documents the formulas and the lookup tables so you can verify any number on the site by hand — and so the rules live in one editable place, not scattered through templates.\nSession conditions Two ambient readings logged at roast start — not used in any formula today, but tracked because they explain session-to-session variation in FC timing:\nField Effect ambient_f Higher temp → earlier FC; compounding heat in back-to-back sessions ambient_rh Higher RH → beans carry more moisture → longer drying phase → FC comes later The two effects can partially cancel: a hot, humid day (91°F / 65% RH) will pull FC earlier via heat but push it later via moisture. Once you have 10+ roasts logged, compare FC times against these readings to see which variable dominates on your machine in your space.\nRaw inputs These are logged per roast. Everything in the \u0026ldquo;Measured\u0026rdquo; panel on a roast page is derived from them — no double-entry, no drift.\nField What it is When to log green_weight_g Weighed grams before the roast Pre-roast roasted_weight_g Weighed grams after cooling Post-roast time_to_fc (mm:ss) Elapsed time from Start to the first audible pop At first crack total_time (mm:ss) Elapsed time from Start to when you hit Cool At drop The two ambient readings above (ambient_f, ambient_rh) are also raw inputs, but they don\u0026rsquo;t feed any formula yet — they\u0026rsquo;re recorded for pattern-finding once batch history grows.\nDerived values The single-page panel computes these automatically; you can match them with a calculator.\nQuantityFormulaWhy it matters Weight loss % (green_weight_g − roasted_weight_g) / green_weight_g × 100 Rough roast-development proxy. Heavier loss = more developed. Development time total_time − time_to_fc How long the bean spent post-first-crack. Drives sweetness and body. DTR (development time ratio) % development_time / total_time × 100 Share of the roast in development. Genuine Origin targets 16–20%; \u0026gt; ~24% risks baked / flat cup (Rao). Worked example For a roast with green_weight_g: 227, roasted_weight_g: 193, time_to_fc: 10:30, total_time: 12:45:\nWeight loss: (227 − 193) / 227 × 100 = 15.0% Development time: 12:45 − 10:30 = 2:15 DTR: 135s / 765s × 100 = 17.6% Resting / ready-to-drink window Fresh-roasted coffee needs to degas before it tastes right. CO₂ trapped in the beans interferes with extraction; brew too soon and the cup tastes flat or sour and espresso channels. Darker roasts shed CO₂ faster than light roasts (longer pyrolysis = more porous structure).\nThe rest_days ranges in the reference table are tuned for a drum roaster brewed by filter methods, which is what this site logs. Two assumptions baked into the numbers:\nDrum vs air. Scott Rao notes classic-drum roasts don\u0026rsquo;t benefit from more than 1–2 days of rest unless underdeveloped; air roasters (Loring, IKAWA) often need 1–4 weeks. The Behmor is a drum roaster. Filter vs espresso. Filter brewing vents CO₂ freely, so short rests work — vendor \u0026ldquo;12–24hr\u0026rdquo; minimums on the bag are real. Espresso runs at 9 bar with nowhere for gas to escape and channels without more rest. For espresso, add ~5–7 days to the upper bound of each range. The single-page panel compares today against the level\u0026rsquo;s window and shows a status pill:\nResting — under the lower bound; CO₂ may still fight extraction Ready — in the peak window; brew and rate Past peak — drinkable for weeks more, but aromatics are fading (faster for dark/oily — Rao recommends drinking those within days) Target level reference These ranges drive the \u0026ldquo;Expected\u0026rdquo; column in the panel. They are industry-consensus starting points and apply to any roaster — calibrate against your own batches over time.\nLevel Weight loss % DTR % Rest (days) Drop trigger Flavor profile light 12–14% 16–19% 1–10 15–45s after FC starts, before any oil sheen Bright acidity, floral/fruit forward, origin-distinct light-medium 13–15% 17–20% 1–7 30–60s into development, before FC fully winds down Balanced acidity, sweetness emerging, origin still legible medium 14–16% 18–22% 1–5 End of FC through 1:30–2:30 of development, before SC Caramel / cocoa, balanced body, low acidity medium-dark 16–18% 19–23% 1–4 First few snaps of second crack Roasty notes appearing, body forward, less origin character dark 18–20% 20–24% 0–2 Rolling SC, before oil floods the surface Smoke / spice / bitter, body dominant Machine-specific guidance This page is intentionally machine-agnostic. For machine-specific guidance — power profiles, weight-setting program times, preheat sequences, safety features — see the page for your roaster:\nBehmor 2000 AB Plus — current machine Sources Genuine Origin — Roast Profiles 101 (DTR 16–20% target across origins) Scott Rao — The Coffee Roaster\u0026rsquo;s Companion (DTR \u0026gt; ~24% risks baking) Sweet Maria\u0026rsquo;s — roast color and loss percentage guide Green Coffee Collective — Types of Coffee Roasts Scott Rao — Resting Roasts: Is Fresher Better? (rest_days; drum vs air roaster differences) Caveats Most published targets — the 8–9 min FC benchmark, DTR ranges, the temperature thresholds — come from drum roasters with bean-probe feedback. Probeless home roasters like the Behmor run a fixed time/power program and are roasted by ear. Treat the numbers as the shape you\u0026rsquo;re aiming for, not setpoints to dial in directly. The reference values on this page get more honest the more of your own batches you log.\n","date":"June 13, 2026","permalink":"/roasts/methods/","summary":"How the numbers on each roast page are calculated, and the lookup tables that drive the expected-range panel.","title":"Methods \u0026 reference","type":"roasts"},{"content":" AI-assisted: I developed the thesis, provided the source materials (the Mikado Method Chapter 1 PDF and a spec-driven development reference), and reviewed the final post for accuracy. Claude helped draft and structure the content.\nThere is a Japanese game called Mikado, which most people know as pick-up sticks. You drop a bundle of sticks on the table and then try to remove them one at a time without disturbing the others. The one you want is always buried under three more. Every large codebase works exactly the same way.\nThe Mikado Method , a book by Ola Ellnestam and Daniel Brolund published in 2014, takes that metaphor seriously. It is a structured technique for making significant changes to complex software without breaking the build at any point along the way. I have found it to be one of the clearest thinking about how software change actually works. And recently I have started to believe it was inadvertently designed for a world where AI writes the code.\nWhat the Mikado Method actually is The method has four primitives.\nSet a goal. Write down what you want the system to do or be when you are finished. Be concrete. \u0026ldquo;Admin services are in a separate package deployable without customer services\u0026rdquo; is a goal. \u0026ldquo;Improve the architecture\u0026rdquo; is not.\nExperiment naively. Try to implement the goal right now, without analyzing all the consequences first. Make the change. Run the compiler. Run the tests. See what breaks.\nVisualize. Whatever breaks is a prerequisite: something that must be true before your goal is achievable. Write it down as a node in a graph, with an arrow pointing toward the goal it unblocks. The graph is the artifact. It is the only thing that survives the next step.\nUndo. Revert every breaking change. Return to the last known working state. Start over on a prerequisite node, not the goal.\nYou repeat this loop for each prerequisite, and each prerequisite\u0026rsquo;s prerequisites, until you reach leaf nodes that can be implemented cleanly without breaking anything. Then you work back up the tree, committing at each step, always keeping the codebase green.\nThe authors are direct about one thing that surprises new practitioners: the undo step is not waste. The code you wrote and reverted taught you something. The graph holds that knowledge. Nothing was lost except the broken code, which was never going to ship anyway.\nWhy it works, and where the friction lives For a human developer working on a brownfield system, the Mikado Method delivers on its promises. The codebase never enters a long-lived broken state. There are no \u0026ldquo;refactoring branches\u0026rdquo; that accumulate months of divergence and produce a painful merge at the end. Each commit is small, green, and shippable. Stakeholders can watch the graph shrink as prerequisites get checked off, which is far better than watching nothing happen for three weeks and then getting a big bang release.\nThe method also surfaces dependencies empirically rather than through analysis. Instead of spending hours reading code trying to predict what will break, you try the change and let the compiler and tests tell you. The naive approach is faster than the analytical one.\nBut I will be honest about the friction, because the method is not free of it.\nThe revert cycle has a cost that the book somewhat underplays. Writing code you know you are about to delete feels bad, even when you understand intellectually that you are learning, not wasting time. There is an emotional overhead to undoing your own work repeatedly. For a developer who is already under pressure, the discipline required to revert rather than keep patching is genuinely hard to maintain.\nThere is also ceremony. Maintaining the graph, deciding which node to work next, keeping the prerequisites from becoming stale as the codebase evolves around you: these are real coordination costs, especially for a solo developer on a tight deadline. The method was designed for teams. For one person it sometimes feels like a lot of scaffolding.\nThe result is a technique that produces excellent outcomes but asks real effort from the human doing it. That tradeoff made sense in 2014. It looks different now.\nThe AI coding agent problem Unconstrained AI coding agents have a characteristic failure mode that I have seen described many ways, but the underlying shape is always the same.\nYou give an agent a goal. It charges at the goal directly. It makes changes across a wide surface area of the codebase simultaneously, because it has read everything and sees all the connections at once. The code it produces compiles, passes a superficial check, and looks plausible. But the failure mode that hurts in practice is subtler than \u0026ldquo;wrong code.\u0026rdquo; Individual diffs look correct in isolation; the system as a whole has a latent bug no single commit would reveal. An agent widens a type across many files and every callsite\u0026rsquo;s tests still pass, but one downstream caller compares strictly where the producer normalizes. Each file is green. The integration is broken. You discover this later through a test you did not have, or through production behavior that looks nothing like the spec.\nThe agent has no intuition about working state. It does not feel the pain of a broken build. It does not develop a reflexive preference for keeping things green. It is optimistic by construction: it produces a complete-looking answer rather than surfacing what it does not know.\nThe longer an agent session runs, the worse this gets. Context drift is well-documented: in long-running sessions, agents lose track of earlier constraints and decisions, gradually producing output that diverges from the original intent. The agent at the end of a two-hour session is operating with a degraded model of the codebase compared to the agent at the start.\nThe fixes people reach for tend to be external: code review, CI gates, smaller prompts, more explicit instructions. These help but they do not address the structural problem, which is that agents need a method, not just guardrails.\nWhere Mikado and AI overlap The Mikado Method maps almost perfectly onto the agent workflow problem.\nMikado concept Agent equivalent Mikado Goal The spec or user story given to the agent Naive implementation Let the agent attempt the change without scaffolding Errors surface prerequisites Test and compile failures reveal what the agent needs first Visualize the graph Human captures failures as structured prerequisite tasks Undo Discard the agent\u0026rsquo;s branch; cost is near zero Work the leaf nodes first Give the agent atomic, prereq-free sub-tasks one at a time The critical inversion is in the undo step.\nFor a human developer, reverting code is emotionally and temporally costly. You wrote that code. It took time. Deleting it requires discipline and a certain tolerance for feeling like you went backward. This is the friction that makes Mikado hard to sustain under pressure.\nFor an AI agent, discarding a branch costs nothing. The agent has no attachment to what it wrote. The tokens were cheap. The wall-clock time was seconds. There is no emotional overhead to reverting. There is no sunk cost fallacy to fight.\nOne clarification worth making. This applies to the exploration phase specifically. The naive experiment is throwaway. The leaf implementations — the actual commits that build up the goal — are kept and shipped. The inversion is not \u0026ldquo;agents write throwaway code.\u0026rdquo; It is that the part humans used to pay for emotionally and temporally (exploring, learning, reverting) is now nearly free, while the graph-building the method always relied on becomes the human\u0026rsquo;s primary contribution.\nThis completely inverts the cost structure of the method. The part that was the bottleneck for humans, doing the implementation work, is now essentially free. The part that was never the bottleneck, building and maintaining the prerequisite graph, is now the critical human contribution. The human\u0026rsquo;s job is not to write the code. It is to run the graph.\nSeveral other ideas arrive at the same place I find it striking how many independent threads in software research converge on this structure.\nReAct , a 2022 paper from Google Research, showed that AI agents perform significantly better when they interleave explicit reasoning traces with actions rather than acting directly. The Think-Act-Observe loop they describe is structurally identical to Mikado\u0026rsquo;s experiment-visualize-undo cycle. The Mikado Graph is what you get when you externalize and persist those reasoning traces across many iterations.\nModern agent orchestration frameworks have independently landed on directed acyclic graphs (DAGs) to represent task dependencies. In Plan-and-Execute frameworks, a planner produces a DAG of sub-tasks, and executors work the leaf nodes first. That is the Mikado graph. The frameworks are reinventing it from the AI side; Mikado invented it from the human side ten years earlier.\nContext drift research consistently identifies two root causes: vague scope and long session duration. Mikado addresses both directly. Each agent session is scoped to a single prerequisite node, which is as constrained as a task definition can be. Each session starts from a fresh context: the agent sees the graph, the leaf\u0026rsquo;s definition, and nothing else. There is no accumulated conversation history to drift from. (Runtime state — caches, container images, already-applied database migrations — does not reset when the session does, and Mikado does not address that side of \u0026ldquo;clean state.\u0026rdquo; More on that in the limits section.)\nUncle Bob\u0026rsquo;s writing on TDD cycles explicitly recommends reverting rather than debugging when a new test does not pass quickly. \u0026ldquo;Backtrack and delete tests until reaching a point from which you can take a different path.\u0026rdquo; Mikado formalizes this instinct into a methodology for large-scale change. TDD validates the revert discipline at the micro level; Mikado applies it at the architectural level.\nGitHub\u0026rsquo;s Agentic CI writing from 2025 describes the emerging pattern where agents produce PRs rather than direct commits. Changes are proposed, tested, and merged only when clean. The Mikado workflow generates exactly these kinds of tasks: small, scoped, green-or-nothing.\nWhat this looks like in practice The workflow I have been thinking through goes like this.\nWrite the Mikado Goal as a concrete spec: a description of the target behavior that can be verified. This is the one input that requires careful human thought. Everything else flows from it.\nIf you already have a plan or design document, feed it to the skill as the starting hypothesis. The naive experiment then becomes a validation pass: prerequisites your plan anticipated get confirmed or refuted, and new ones the plan missed surface as failures. Starting with a plan does not skip the experiment — it gives you something concrete to test against.\nPoint an agent at the goal with no scaffolding. Tell it to implement the goal directly. Do not help it avoid the obstacles. You want to find the obstacles.\nCollect the failures, not the code. Compile errors, failing tests, runtime exceptions: these are the data you are after. The code the agent wrote is almost certainly going to be reverted.\nBuild the prerequisite graph from the failures. This is the human judgment step, and it is the most important one. Deciding which failures share a root cause, which prerequisites are truly independent, and which order makes sense requires understanding the system. This is not a step you can delegate to the agent.\nAssign a single leaf-node prerequisite to the agent. One task, clean scope. Whether that runs in the current session or a fresh one depends on the leaf\u0026rsquo;s size and surface area — small, self-contained changes often stay inline; anything that crosses subsystem boundaries or introduces a new abstraction delegates to a fresh context.\n\u0026ldquo;Single leaf\u0026rdquo; is directionally right but not always literal. Mechanical refactors — type widenings, renames, database column migrations — naturally span many files because intermediate states do not compile. Judge leaf size by cognitive load, not by line count or file count. Ten files of the same mechanical edit repeated is one leaf. Two files with subtle coupling and a new abstraction introduced between them is two leaves.\nCommit only when the session\u0026rsquo;s sub-task passes all tests cleanly. If it does not pass, revert and refine the prerequisite definition. The graph is a living document; new prerequisites emerge as you work. Some will surface during the naive experiment; others only appear once a leaf is underway. When a leaf is halfway done and you realize it depends on a sibling prerequisite you had considered independent, stop the leaf, add the dependency to the graph, and re-pick. Resist the urge to fold the sibling\u0026rsquo;s work into the current leaf to finish faster. That is how scope leaks — the quiet failure mode of the method — happen. A leaf that solves its surface symptom by reaching into a sibling\u0026rsquo;s scope creates latent gaps that look fine at review time and break in production.\nRepeat until the Mikado Goal is met.\nThe description above frames the agent as the one implementing each leaf. The method also works with the roles inverted. If you have deep domain knowledge in the system — the kind that comes from years of working in the codebase, or from being the person who designed the feature — you may not want to hand implementation to an agent at all. In that case the agent becomes a different kind of collaborator: it runs the naive experiments, builds and maintains the prerequisite graph, derives the testing plan, and verifies each leaf after you commit it. You implement. At each step the agent hands you a single, clearly scoped task with verified acceptance criteria, marks it done when it passes, and picks the next one. The undo cost is still near zero — if your implementation reveals that a leaf needs to be split, the agent records the sub-prerequisites and re-queues. The discipline the method imposes is unchanged. Who writes the code is a configuration choice, not a constraint.\nThe output of this process is a codebase that was never broken, a commit history that tells a coherent story, and a graph that documents the dependency structure you discovered along the way.\nMaking it real: tooling the method The method described above is a workflow. In practice, it starts to compound only once the workflow is codified as tooling.\nIn my own setup, Mikado lives as a stack of composable skills (available on GitHub ): one to kick off a goal and run the naive experiment, one to drive a single leaf to completion — inline or in a fresh context depending on the leaf\u0026rsquo;s scope, one to assemble the pull request from the graph and the commits it produced. A project-specific wrapper sits on top and codifies conventions — known flaky tests, narrow test commands, codegen rules, commit message formats — so every subagent receives them identically without a human remembering to include them in the prompt. The skills compose via a loop orchestrator so the full run is effectively loop \u0026lt;mikado-leaf\u0026gt; until the graph is empty.\nThe distinction matters because fresh-session-per-leaf is incompatible with tribal knowledge. A rule like \u0026ldquo;never run the full test suite, only the affected test class\u0026rdquo; that lives in your head will be included in the first leaf\u0026rsquo;s prompt and forgotten by the fifth. The same rule embedded in the skill\u0026rsquo;s operating rules is applied identically every time. The move from \u0026ldquo;method I follow\u0026rdquo; to \u0026ldquo;method my tools enforce\u0026rdquo; is where the method\u0026rsquo;s consistency claims start actually holding.\nThe right amount of tooling depends on the project. A small or new codebase may need only the four core skills. A large monorepo benefits from a project wrapper that catalogs its known flakes, its narrow test commands, its codegen pipeline, and its commit conventions. Either way, getting out of the business of remembering these things is what makes the method scale past one or two goals.\nThe limits The human still builds and maintains the graph. This is not a step you can automate away. Reading the agent\u0026rsquo;s failures and understanding what they mean requires familiarity with the codebase and judgment about the domain. The method reduces the implementation burden on the human dramatically; it does not reduce the thinking burden.\nThe method also requires fast feedback. Mikado was designed for systems with compilers and test suites. The faster the feedback loop, the cheaper each experiment is. If your test suite takes twenty minutes to run, the per-experiment cost goes back up. This is not a new constraint; it is the same one Mikado has always had.\nAnd the graph is a hypothesis, not a plan. Leaf nodes that look independent often reveal new dependencies when you actually implement them. The tree grows as you work it. Treat the graph as something you are discovering, not something you designed upfront. The moment you mistake it for a waterfall plan, you will start ignoring what the failures are trying to tell you.\nGreen leaves are not a green system. Each leaf\u0026rsquo;s tests verify the leaf. Integration bugs that span multiple leaves — a type widening that changes behavior at callsites which were individually green but combine incorrectly, a data format change that different callers parse with slightly different rules — still slip through. The method produces commits that are small, bisectable, and reviewable, which makes integration-level review tractable. It does not make that review unnecessary. The human who reviews the full diff is still doing work the method cannot replace.\nClean state means the runtime, not just the session. A fresh agent session starts with no conversation history. The runtime environment — caches, container images, hot-reload classloaders, already-applied database migrations — does not reset with it. A Mikado run against a live system can be sabotaged by a stale Redis entry, a partially-rolled deployment, or a data migration that ran against the wrong codebase version. These are not context drift in the LLM sense; they are state drift in the operational sense, and they need to be addressed by the same discipline that keeps any production system consistent. The method assumes you are doing that work separately.\nThe graph structure tells you how to ship. A Mikado graph with genuinely independent prerequisites can ship as independent pull requests reviewed in parallel by the right subject-matter reviewers. A linear chain of prerequisites ships as a stacked PR set or a single PR. If you default to \u0026ldquo;one goal, one PR,\u0026rdquo; you may leave the graph\u0026rsquo;s natural shape on the floor and create a larger review surface than the work required. The graph is an asset for the delivery flow, not just the implementation flow.\nConclusion The Mikado Method was not designed for AI agents. It was designed for human developers working in brownfield systems that had no tests and no documentation, trying to make large changes without destroying the codebase in the process. The authors built it from years of experience doing exactly that work.\nBut the method\u0026rsquo;s core insight turns out to be more durable than its original context. Exploration is learning. Reversion is not waste. The graph is the product. These are true regardless of who is doing the implementation.\nWhen a human is doing the implementation, maintaining those principles requires discipline and emotional resilience. The undo cycle costs something real. When an agent is doing the implementation, the undo cycle costs nothing. The entire friction model shifts. The graph, which was always the most valuable thing the method produced, becomes even more valuable because the human\u0026rsquo;s time is now freed entirely to build it.\nThe developers who will get the most out of AI coding agents are not the ones who can write the best prompts. They are the ones who can think clearly about prerequisites: what needs to be true before the thing you want is possible, and in what order those things should happen. That has always been what the Mikado Method was teaching.\n","date":"March 1, 2026","permalink":"/posts/mikado-method-ai-agents/","summary":"A 2014 technique for navigating legacy codebases turns out to be a near-perfect framework for governing AI coding agents, once you account for the fact that AI-generated code is essentially free.","title":"The Mikado Method in the Age of AI Agents","type":"posts"},{"content":" AI-assisted: I researched the topics covered here, developed the outline and key arguments, and used Claude to help draft and compile the final post. All content was reviewed for accuracy, and quotes and statistics are attributed with source links.\nContent sourced from MuleSoft, Boomi, WSO2, Solace, Confluent, CNCF, Capital One, eBay, and other industry leaders. All quotes and statistics are attributed with source links.\nExecutive Summary In the middleware and enterprise integration industry, schemas and message standards are not merely technical details. They are foundational contracts that enable systems to exchange business data reliably at scale. This document explores why leading integration platforms like MuleSoft, Boomi, WSO2, and others are investing heavily in open standards, and why products that understand, respect, and utilize schema definitions deliver superior business value.\nThe core business problem: Your order management system needs to send purchase orders to your ERP. Your ERP needs to update your warehouse system. Your warehouse system needs to notify your shipping provider. Each system was built by different vendors with different assumptions about data formats. Without standardized schemas, every integration becomes a custom project. Expensive, fragile, and difficult to maintain.\nThe solution: Schema standards like OpenAPI, AsyncAPI, CloudEvents, and Schema Registry provide common \u0026ldquo;contracts\u0026rdquo; that define how systems exchange data. With them, organizations gain interoperability, governance, and the ability to evolve systems without breaking downstream business processes.\n1. Why Schemas Matter in Enterprise Integration 1.1 The Problem of Integration Chaos Consider a typical business scenario: your order management system needs to send purchase order data to your ERP, your warehouse management system, and your supplier\u0026rsquo;s fulfillment platform. Each system was built by a different vendor, at a different time, with different assumptions about how a \u0026ldquo;purchase order\u0026rdquo; should be structured.\nThe problem? Every system describes the same business data differently. The lack of a common way of describing business transactions means developers are constantly re-learning how to exchange data with each new system. This also limits the potential for reusable tools and infrastructure to aid the delivery of business data across environments.\n\u0026ldquo;The portability and productivity that can be achieved from event data is hindered overall.\u0026rdquo;\n— CloudEvents Specification, CNCF (source )\nThis challenge is not abstract. According to Airbyte\u0026rsquo;s analysis on schema evolution , enterprises average one schema modification every 3.03 days across typical systems. Without proper management, these changes break integration workflows, corrupt reports and dashboards, and derail downstream business processes.\nReal-world example: A field called employeeID exists in your HR system. Your payroll system expects it. One day, the HR team removes that field during a system upgrade. The payroll integration, unaware of this change, fails silently or crashes, and suddenly employees aren\u0026rsquo;t getting paid correctly.\n1.2 Schemas as Contracts A schema defines the structure of business data exchanged between systems. It specifies allowed data types, their format, and relationships. Think of a schema as a blueprint that describes:\nWhat fields are included in a purchase order, invoice, or customer record What data type each field contains (text, number, date, etc.) Which fields are required vs. optional How fields relate to each other \u0026ldquo;Schema Registry allows for schema evolution and provides compatibility checks to ensure that the contract between producers and consumers is not broken. This allows producers and consumers to update independently and evolve their schemas independently, with assurances that they can read new and legacy data.\u0026rdquo;\n— Confluent Documentation (source )\nIn business terms: when your order system sends data to your warehouse system, both systems agree on what an \u0026ldquo;order\u0026rdquo; looks like. If the order system changes its format, the schema contract ensures the warehouse system can still understand the data, or alerts you that a breaking change is about to occur.\n1.3 The Business Impact Integration platform vendors recognize this reality. The shift toward standardized schemas is accelerating because enterprises need systems that can reliably communicate across organizational boundaries, including partners, suppliers, and customers.\n\u0026ldquo;Event-driven integration is being adopted as the go-to architecture and so standards like CloudEvents are essential to increasing interconnectivity, inside and across enterprises.\u0026rdquo;\n— Solace (source )\nProducts that fail to support schema standards force customers into manual workarounds, increased development complexity, and a lack of consistency across their integration landscape. As noted in the MuleSoft blog on AsyncAPI and CloudEvents , relying solely on standard connectors can lead to these challenges when native support is limited.\nThe bottom line: Without schema standards, every integration becomes a custom project. With them, integrations become predictable, maintainable, and scalable.\n2. The Major Schema Standards Landscape The integration industry has developed several complementary standards to address different communication patterns. Understanding when to use each is key to building a coherent integration strategy.\nStandard Primary Use Case Think of It As\u0026hellip; OpenAPI Request/response APIs (REST) Defining how systems ask for and receive data AsyncAPI Asynchronous messaging Defining how systems send notifications and updates CloudEvents Message envelope format A standard \u0026ldquo;shipping label\u0026rdquo; for any message Schema Registry Schema versioning \u0026amp; governance A central catalog of all data formats OTel Semantic Conventions Observability data naming Consistent names for monitoring/tracing data EDI (X12/EDIFACT) B2B document exchange Standardized business documents for trading partners 2.1 OpenAPI: The Foundation for Request/Response APIs When one system needs to request data from another, like querying a customer\u0026rsquo;s order history or submitting a new invoice, this typically happens through a REST API. The OpenAPI Specification (OAS) is the widely adopted standard for describing these APIs.\nAccording to the OpenAPI Initiative , OpenAPI has become the de facto standard for REST APIs, providing a human and machine-readable structure that describes:\nWhat endpoints are available (e.g., /orders, /customers/{id}) What data you send in a request What data you receive in a response How authentication works \u0026ldquo;This allows people to understand how an API works, how a sequence of APIs work together, generate client code, create tests, apply design standards, and much, much more.\u0026rdquo;\n— OpenAPI Initiative (source )\nIndustry Adoption: According to Wikipedia\u0026rsquo;s OpenAPI article , founding members of the OpenAPI Initiative in 2015 include Google, IBM, Microsoft, PayPal, Capital One, 3Scale, Apigee, Intuit, and Restlet. MuleSoft joined in 2017 after developing RAML.\nBusiness Value: When your team needs to integrate with a partner\u0026rsquo;s system, an OpenAPI specification tells your developers exactly what to expect. No guesswork, no back-and-forth emails asking \u0026ldquo;what format should the date be in?\u0026rdquo; Per Boomi\u0026rsquo;s documentation , the platform automatically generates OpenAPI specification files for each deployed REST API.\nThe 2024 State of SaaS APIs report found that OpenAPI and Swagger are the most supported API specification formats, though noted that \u0026ldquo;many developers see the necessary step of making API specifications as an unnecessary formality,\u0026rdquo; explaining why many APIs remain poorly designed.\nKey Benefits of OpenAPI Benefit Business Value Design-First Development Define the integration contract before writing code, catching misunderstandings early Auto-Documentation Generate interactive documentation so partners can self-serve Code Generation Automatically create integration code, reducing development time Contract Testing Validate that implementations match specifications before go-live Standardization Common language across teams, vendors, and partners 2.2 AsyncAPI: Defining Asynchronous Business Communications Not all business processes work as simple request/response. Consider these scenarios:\nOrder placed → Notify warehouse, update inventory, send confirmation email (one action triggers multiple downstream processes) Inventory below threshold → Alert purchasing department (system monitors and notifies when conditions are met) Payment received → Update accounts receivable, trigger fulfillment (one system informs others that something happened) These patterns require asynchronous communication, where the sending system doesn\u0026rsquo;t wait for an immediate response. AsyncAPI is the specification for describing these types of integrations.\nAccording to the AsyncAPI Initiative , AsyncAPI is protocol-agnostic, supporting various messaging technologies (AMQP, MQTT, WebSockets, Kafka, and more).\n\u0026ldquo;Evented APIs arose as a way to bring a similar level of governance to the EVENTful world that already existed in the RESTful world, and to make the experience of the API consumer similar as well.\u0026rdquo;\n— MuleSoft Blog on Evented APIs (source )\nThe specification separates integrations into three layers, as described in MuleSoft\u0026rsquo;s AsyncAPI documentation :\nMessages: The business data being exchanged (e.g., order details, inventory update) Channels: Where the message is sent (e.g., \u0026ldquo;orders\u0026rdquo; queue, \u0026ldquo;inventory-updates\u0026rdquo; topic) Transport: The underlying technology (Kafka, RabbitMQ, etc.) Real-World Adoption:\nOrganization Use Case Source eBay Standardized notification system for marketplace updates eBay Tech Blog Capital One Contract-first development for internal system communication Capital One Tech TransferGo Blueprint for payment processing workflows AsyncAPI Case Studies Slack Documenting data streaming APIs for partners AsyncAPI Case Studies MuleSoft Implementation: Per MuleSoft\u0026rsquo;s release notes , Anypoint Platform supports AsyncAPI 2.6 in Anypoint Code Builder for designing API specs and Avro schemas, API Designer, and API Experience Hub.\n\u0026ldquo;AsyncAPI addresses the need for a unified, open source, protocol-agnostic asynchronous specification that is both human-readable and machine-readable, while also being backed by a diverse and rich tooling ecosystem.\u0026rdquo;\n— eBay Tech Blog (source )\n2.3 CloudEvents: A Standard Envelope for Business Messages When systems exchange business data, the message itself (the order, the invoice, the status update) is only part of the picture. Receiving systems also need to know:\nWhere did this message come from? What type of message is this? When was it created? How should I interpret the payload? CloudEvents is a CNCF-graduated specification that standardizes this \u0026ldquo;envelope\u0026rdquo; information. Think of it like a shipping label that accompanies every package. Regardless of what\u0026rsquo;s inside, the label follows a standard format that any logistics system can read.\nAccording to the CNCF announcement , CloudEvents was approved as a graduated project on January 25, 2024.\n\u0026ldquo;It\u0026rsquo;s a simple data model with a focus on describing the context of events that works with numerous popular protocols and data encodings in a uniform way, without trying to invent new protocol features or getting in the way of the protocol features that exist.\u0026rdquo;\n— Microsoft Principal Architect Clemens Vasters (source )\nEnterprise Scale:\nPer The New Stack\u0026rsquo;s coverage :\nMicrosoft uses CloudEvents across Microsoft 365 and Azure cloud services, processing billions of messages each day SAP uses CloudEvents as a foundation across SAP Event Broker, SAP BTP Kyma runtime, SAP Event Mesh, and SAP Integration Suite \u0026ldquo;The CloudEvents specification has been used as a foundation in SAP\u0026rsquo;s Business Technology Platform for alignment across applications and platform services\u0026hellip; Adhering to the open CloudEvents specification provides interoperability across messaging protocols and its extensibility simplifies customer adoption.\u0026rdquo;\n— SAP, via CNCF Graduation Announcement (source )\nWhy This Matters for Integration Products: CloudEvents allows different systems, potentially from different vendors, to exchange messages without custom translation of the envelope metadata. The business payload can be anything (an order, an alert, a status update), but the routing and handling information follows a universal standard.\n\u0026ldquo;It\u0026rsquo;s not creating a whole new set of middleware, and we were very conscious not to do that. We wanted people to slowly adapt to CloudEvents without having to sort of reinvent their entire middleware architecture.\u0026rdquo;\n— Doug Davis, Co-chair of CloudEvents and CNCF Serverless WG (source )\nCloudEvents Core Attributes Per the CloudEvents Specification :\nAttribute Required Business Purpose id Yes Unique identifier. Enables deduplication and tracking source Yes Which system sent this (e.g., \u0026ldquo;orders.acme.com\u0026rdquo;) specversion Yes CloudEvents version for compatibility type Yes What kind of message (e.g., \u0026ldquo;com.acme.order.created\u0026rdquo;) datacontenttype No Format of the payload (e.g., \u0026ldquo;application/json\u0026rdquo;) dataschema No Link to the schema defining the payload structure subject No Specific resource this relates to (e.g., order ID) time No When this occurred. Critical for sequencing 2.4 Schema Registry: Centralized Catalog of Data Formats As your organization builds more integrations, you quickly accumulate dozens (or hundreds) of different data formats: order schemas, customer schemas, inventory schemas, each potentially with multiple versions. How do you keep track of them all? How do you ensure that when System A sends data, System B knows how to interpret it?\nA Schema Registry is a centralized catalog that stores, versions, and validates all your data schemas. It serves as the \u0026ldquo;single source of truth\u0026rdquo; for what your business data looks like.\n\u0026ldquo;Schema Registry provides several benefits, including data validation, compatibility checking, versioning, and evolution. It also simplifies the development and maintenance of data pipelines and reduces the risk of data compatibility issues, data corruption, and data loss.\u0026rdquo;\n— Confluent Documentation (source )\nWhy It Matters: Per Confluent\u0026rsquo;s Schema Registry course :\n\u0026ldquo;In any non-trivial system, the Schema Registry is indispensable. As applications grow, new consumers emerge, and schemas inevitably evolve. The Schema Registry provides centralized schema management – all producers and consumers share a common understanding of message formats.\u0026rdquo;\nHow It Works in Practice:\nAccording to Confluent\u0026rsquo;s tutorial :\nRegistration: When your order system is ready to send order data, it first registers the order schema with the registry ID Assignment: The registry assigns a unique schema ID and returns it Message Tagging: Your order system includes this schema ID with every order message it sends Lookup: When your warehouse system receives an order message, it uses the schema ID to look up the exact format Validation: The warehouse system deserializes the message using the correct schema, guaranteed to match Business Benefit: Your order system and warehouse system can be updated independently. As long as schema changes follow compatibility rules, neither system breaks when the other is updated.\nSupported Formats: Per Confluent\u0026rsquo;s Schema Registry Overview :\nFormat Characteristics Best For Avro Compact binary format, built-in schema evolution High-volume transaction processing JSON Schema Human-readable, widely understood Web integrations, debugging, partner APIs Protobuf Efficient, language-neutral, gRPC native Cross-platform microservices \u0026ldquo;Schema Registry supports multiple formats at the same time. For example, you can have Avro schemas in one subject and Protobuf schemas in another.\u0026rdquo;\n— Confluent Documentation (source )\n2.5 Related Standard: OpenTelemetry Semantic Conventions While the standards above define the structure of business data exchanged between systems, there\u0026rsquo;s a complementary standard worth understanding: OpenTelemetry Semantic Conventions. These conventions define consistent naming for observability data, the telemetry (traces, metrics, logs) that monitors how your integrations perform.\nWhat It Does: OpenTelemetry Semantic Conventions specify common names for different kinds of operations and data across your technology stack. According to the OpenTelemetry documentation :\n\u0026ldquo;The benefit of using Semantic Conventions is in following a common naming scheme that can be standardized across a codebase, libraries, and platforms.\u0026rdquo;\nWhy This Matters for Integration: When your order management system sends data to your warehouse system, you use OpenAPI or AsyncAPI to define that data contract. But when you want to monitor that integration (track latency, identify errors, correlate logs across services) you need consistent naming for your observability data. That\u0026rsquo;s where OTel Semantic Conventions come in.\nThink of it this way:\nStandard Answers the Question OpenAPI / AsyncAPI \u0026ldquo;What does the order data look like?\u0026rdquo; CloudEvents \u0026ldquo;What metadata accompanies each message?\u0026rdquo; OTel Semantic Conventions \u0026ldquo;What do we call the fields when we log or trace this operation?\u0026rdquo; Industry Convergence: In April 2023, Elastic and OpenTelemetry announced plans to converge the Elastic Common Schema (ECS) with OTel Semantic Conventions , with the goal of creating a single unified schema:\n\u0026ldquo;The goal is to achieve convergence of ECS and OTel Semantic Conventions into a single open schema that is maintained by OpenTelemetry, so that OpenTelemetry Semantic Conventions truly is a successor of the Elastic Common Schema.\u0026rdquo;\n— OpenTelemetry Blog (source )\nThis convergence effort brings ECS\u0026rsquo;s mature schema for logs and security events into the OTel ecosystem, making it the emerging standard for all observability naming.\nStability Progress: The conventions are progressively reaching stability:\nHTTP conventions: Declared stable in 2023 (v1.23.0) (source ) Database conventions: Reached stability in 2024 (source ) Messaging conventions: Still in development, covering Kafka, RabbitMQ, AWS SQS/SNS, and more (source ) Integration Platform Adoption:\nMuleSoft now includes native OpenTelemetry support. Per MuleSoft\u0026rsquo;s documentation :\n\u0026ldquo;MuleSoft follows the semantic convention defined by OpenTelemetry that adds one or more exception events to the span\u0026hellip; OpenTelemetry enables Mule runtime engine to provide observability into the behavior of Mule applications.\u0026rdquo;\nBoomi announced early access for OpenTelemetry as part of their 2025 platform releases, allowing native configuration of runtimes to stream real-time telemetry data to preferred third-party monitoring platforms (source ).\nCloudEvents Interoperability: OpenTelemetry includes specific semantic conventions for CloudEvents , enabling correlation between the business messages (CloudEvents format) and the observability traces that track their journey through your systems.\nKey Takeaway: OTel Semantic Conventions are complementary to the data contract standards in this document. You define your API contracts with OpenAPI/AsyncAPI and your message envelopes with CloudEvents, then you use OTel Semantic Conventions to ensure consistent, vendor-neutral monitoring of those integrations. Integration platforms that support both give you governance over your data and visibility into how that data flows.\n3. EDI: The Bedrock of B2B Data Exchange 3.1 Understanding EDI Standards Electronic Data Interchange (EDI) is the computer-to-computer exchange of business documents in a standard electronic format between business partners. According to Orderful\u0026rsquo;s X12 overview , EDI standards have been around for over 40 years and remain foundational to global commerce.\n\u0026ldquo;X12 is a standard developed by the Accredited Standards Committee (ASC) X12, chartered by the American National Standards Institute (ANSI) in 1979, for interindustry electronic exchange of business transactions.\u0026rdquo;\n— Introduction to B2B EDI (source )\nX12 (ANSI): The primary standard for EDI in North America. Per SEEBURGER\u0026rsquo;s analysis , X12 is used across healthcare (HIPAA compliance), retail, manufacturing, finance, and logistics. It uses three-digit codes (like 850 for Purchase Order, 810 for Invoice).\nEDIFACT (UN/ECE): The international standard used throughout Europe and Asia, developed in 1986. According to EZCom Software , it uses six-letter codes (like ORDERS for Purchase Order, INVOIC for Invoice).\n\u0026ldquo;Today, EDI X12 powers billions of transactions every day across all industries, but especially in retail, manufacturing, healthcare, finance, and logistics. Its standardized formats ensure that critical business documents such as purchase orders, invoices, and shipping notices are exchanged accurately and efficiently between trading partners.\u0026rdquo;\n— EDI2XML (source )\n3.2 X12 vs EDIFACT Comparison Per Zenbridge\u0026rsquo;s analysis :\nAspect X12 EDIFACT Region North America Europe, Asia, International Developed 1979 (ANSI) 1986 (UN/ECE) Document Codes 3-digit (850, 810) 6-letter (ORDERS, INVOIC) Industries Healthcare, Retail, Finance Global Trade, Logistics Governance ASC X12 UN/CEFACT Interchange Header ISA/IEA UNB/UNZ Functional Group GS/GE UNG/UNE Transaction Set ST/SE UNH/UNT 3.3 Why EDI Still Matters for Integration Products Integration platforms must support EDI because it remains the backbone of supply chain and healthcare communications. As noted by GraceBlood :\n\u0026ldquo;While newer technologies like XML, JSON, and APIs are reshaping how data is exchanged, ANSI X12 and EDIFACT remain foundational pillars in EDI—trusted for their robustness, reliability, and widespread adoption.\u0026rdquo;\nProducts like MuleSoft, Boomi, and SEEBURGER provide B2B/EDI capabilities. Per SEEBURGER :\n\u0026ldquo;With SEEBURGER BIS Platform B2B/EDI capabilities and Cloud Integration Services for B2B/EDI, you can convert any EDI data format found on the market.\u0026rdquo;\nKey Insight: An integration platform that cannot handle EDI excludes itself from significant enterprise use cases, particularly in retail, healthcare, logistics, and manufacturing.\nCommon EDI Transaction Types Per CData\u0026rsquo;s X12 EDI Guide :\nX12 Code EDIFACT Document Type 850 ORDERS Purchase Order 810 INVOIC Invoice 856 DESADV Ship Notice/Manifest 855 ORDRSP Purchase Order Acknowledgment 820 REMADV Payment Order/Remittance Advice 997 CONTRL Functional Acknowledgment 4. How Leading Platforms Leverage Schema Standards 4.1 MuleSoft: Full Lifecycle Schema Support MuleSoft announced open beta for AsyncAPI support in Anypoint Platform with a vendor-neutral approach to integrating with various event brokers. Per Solace\u0026rsquo;s coverage of the announcement:\n\u0026ldquo;Traditional iPaaS and integration vendors will hop on the train as customer demand increases – with this announcement MuleSoft joins Boomi, SAP and others in doing so.\u0026rdquo;\nFrom the MuleSoft Blog on Evented APIs :\n\u0026ldquo;AsyncAPI spec and event schemas can be published in a catalog/app store to facilitate discoverability and reuse—addressing one of the event-driven architecture gaps. Once an AsyncAPI is published, mock tests can be created to reflect the utilization of the API, removing the testability gap in event-driven architectures.\u0026rdquo;\nCapabilities:\nAsyncAPI 2.6 support in Anypoint Code Builder (source ) Avro schema design and validation API Experience Hub for discovery Scaffolding generation from AsyncAPI specs (source ) Limitations: As integration architect Giacomo Bartoloni noted in his Medium article :\n\u0026ldquo;While MuleSoft provides robust tools for building event-driven architectures, the limitations in its native AsyncAPI support for certain message brokers necessitate the development of custom connectors to facilitate the usage of CloudEvents format.\u0026rdquo;\n4.2 Boomi: Canonical Data Models Boomi API Management provides central governance for all API connections. Per Boomi\u0026rsquo;s blog on Universal API Management :\n\u0026ldquo;Pre-built connectors include standard data models for 300,000+ unique endpoints, which means developers start with established structures instead of building mappings from scratch. Visual mapping tools let teams define canonical models without writing code, and 200+ million anonymized integration patterns inform machine learning-powered mapping suggestions.\u0026rdquo;\nCanonical Model Approach:\n\u0026ldquo;One Canonical Data Model Reduces Mapping Work: The platform defines standard field names, data types, and structures that all connected applications use. Without this standardization, developers spend time reconciling conflicting definitions where different systems use different naming conventions and formats for the same information.\u0026rdquo;\n— Boomi (source )\nPer Boomi\u0026rsquo;s API Management documentation , the platform automatically generates a WSDL for each deployed SOAP API and an OpenAPI specification file for each deployed REST API.\n4.3 WSO2: Open Standards Focus WSO2 API Manager is an open platform for managing APIs with full lifecycle capabilities. Per WSO2\u0026rsquo;s website :\n\u0026ldquo;WSO2\u0026rsquo;s API management provides the #1 open source, market-leading full lifecycle platform for building, integrating, securing, and exposing AI and digital services as managed APIs in the cloud, on-premises, hybrid architectures, and modern environments like Kubernetes.\u0026rdquo;\nAccording to WSO2\u0026rsquo;s feature documentation , the platform supports:\nAPI-first integration with built-in support for a wide range of protocols (HTTP, AMQP, JMS, Kafka, gRPC) Standards support (OpenAPI, SOAP) Schema validation: \u0026ldquo;Validates APIs payload content against schemas\u0026rdquo; JSON web tokens for consumption by back-end servers Per WSO2\u0026rsquo;s architecture documentation :\n\u0026ldquo;WSO2 Universal Gateway acts as the entry point for an API request made to an API managed by WSO2 API Manager. The Universal Gateway does the JWT token validation by validating the signature, issuer, expiry time, and subscription.\u0026rdquo;\nKey Differentiator: WSO2\u0026rsquo;s open-source approach allows customers to avoid vendor lock-in. Per WSO2 :\n\u0026ldquo;Protect yourself from vendor lock-in, and enjoy complete trust and transparency with our fully open-source software.\u0026rdquo;\n4.4 Solace: Event Portal and AsyncAPI Per Solace\u0026rsquo;s blog on MuleSoft integration :\n\u0026ldquo;Event Portal lets users easily find existing event types or create new event types through an easy set of dialogs. Event Portal also abstracts away AsyncAPI limitations such as ownership, versioning and reusability of each AsyncAPI component (channels, schemas). Finally, Event Portal provides an easier path to AsyncAPI document creation and helps get developers writing MuleSoft Integration Flows faster.\u0026rdquo;\nMuleSoft Integration:\n\u0026ldquo;Creating flows for event-driven integrations can be a real pain, and integrations are frequently way more complicated than they need to be. We know that, as many of us come from writing integration flows ourselves. We\u0026rsquo;ve developed a plugin that lets you export an AsyncAPI specification from Event Portal and generate a MuleSoft Integration Flow, all within the Anypoint Studio IDE.\u0026rdquo;\n— Solace (source )\nPlatform Comparison Matrix Capability MuleSoft Boomi WSO2 Solace OpenAPI Support ✅ Full ✅ Auto-generate ✅ Full N/A AsyncAPI Support ✅ 2.6 (source ) Partial ✅ ✅ Full CloudEvents Custom connector needed (source ) N/A N/A ✅ Native Schema Registry External integration N/A N/A ✅ Built-in (source ) EDI Support ✅ B2B/EDI ✅ B2B/EDI Via integration N/A Open Source No No ✅ Yes (source ) No 5. Schema Evolution: Managing Change Without Breaking Systems 5.1 The Challenge of Change Schemas will never stop changing. New fields get added, types are updated, old columns are dropped. Per Estuary\u0026rsquo;s guide on schema evolution :\n\u0026ldquo;Schema evolution is one of the most complex parts of real-time systems. Without it, pipelines break, dashboards fail, and downstream consumers lose trust.\u0026rdquo;\nFrom Confluent\u0026rsquo;s documentation on schema evolution :\n\u0026ldquo;An important aspect of data management is schema evolution. After the initial schema is defined, applications may need to evolve it over time. When this happens, it\u0026rsquo;s critical for the downstream consumers to be able to handle data encoded with both the old and the new schema seamlessly. This is an area that tends to be overlooked in practice until you run into your first production issues.\u0026rdquo;\n5.2 Compatibility Strategies Per Confluent\u0026rsquo;s Schema Evolution documentation , Schema Registry supports compatibility rules that help manage schema changes gracefully:\nStrategy Description Use When BACKWARD New schema can read old data Upgrading consumers first FORWARD Old schema can read new data Upgrading producers first FULL Both backward and forward Gradual rollouts, safest option NONE No compatibility checking Development/testing only Transitive vs Non-Transitive:\n\u0026ldquo;Transitive compatibility checking is important once you have more than two versions of a schema for a given subject. If compatibility is configured as transitive, then it checks compatibility of a new schema against all previously registered schemas; otherwise, it checks compatibility of a new schema only against the latest schema.\u0026rdquo;\n— Confluent (source )\n5.3 Best Practices for Schema Evolution Per Solace\u0026rsquo;s Schema Registry Best Practices :\nProvide default values for new fields to maintain backward compatibility Never rename fields: \u0026ldquo;Instead of renaming fields, add new fields and deprecate old ones. Renaming breaks compatibility.\u0026rdquo; Make new fields optional to ensure older consumers can ignore them Use semantic versioning: \u0026ldquo;Use semantic versioning for your schemas. This helps track changes systematically.\u0026rdquo; Use union types: \u0026ldquo;Union types allow a field to accept multiple data types, providing flexibility for future changes.\u0026rdquo; Anti-patterns to avoid, per Solace :\n\u0026ldquo;Breaking changes without planning—Avoid removing required fields, changing field types incompatibly, or renaming fields without aliases. Tight coupling to schema versions—Avoid hard-coding schema versions directly in application code.\u0026rdquo;\n5.4 The Outbox Pattern One approach to prevent schema changes from impacting downstream systems is the outbox pattern. Per Decodable\u0026rsquo;s article on schema evolution in CDC pipelines :\n\u0026ldquo;One approach is to use the outbox pattern, hence, not directly exposing data model internals to the outside world to begin with. So instead of configuring a CDC pipeline directly against the database tables backing some application, a contract is defined based on a well-defined public schema.\u0026rdquo;\nHow it works:\nInstead of directly exposing data model internals, define a contract based on a well-defined public schema Data changes are written to a dedicated outbox table following the defined contract Both original table writes and outbox writes are bound to the same database transaction CDC reads from the outbox table, not the internal tables Benefits:\nDecouples internal data models from external contracts Allows internal schema changes without breaking consumers Provides transaction consistency for event publishing 6. The Business Case for Schema-Aware Products 6.1 Avoiding Vendor Lock-In Closed or throttled APIs trap organisations inside proprietary ecosystems. Per research on Vendor Lock-In and Interoperability :\n\u0026ldquo;By prioritizing interoperability, organizations can avoid vendor lock-in, enhance flexibility, and foster innovation, ultimately enabling a more agile and resilient IT infrastructure in the cloud.\u0026rdquo;\nPer ITLawCo\u0026rsquo;s analysis :\n\u0026ldquo;Closed or throttled APIs trap organisations inside proprietary ecosystems. Solution: build abstraction layers or middleware that interface via open standards.\u0026rdquo;\nStrategies for Avoiding Lock-In:\nPer Acceldata\u0026rsquo;s guide on data interoperability :\nStandardization: Adopt industry-standard data formats, protocols, and interfaces Openness: Embrace open standards and APIs to facilitate data exchange and prevent vendor lock-in Metadata management: Establish clear metadata standards to provide context and enable data discovery 6.2 Developer Productivity When working with event-driven architectures, developers need to work with multiple protocols. Per MuleSoft\u0026rsquo;s blog on evented APIs :\n\u0026ldquo;Developer experience: When working with event-driven architectures, developers need to work with multiple protocols (JMS, MQTT, AMQP, Kafka, etc.), which can be challenging.\u0026rdquo;\nCode Generation: Per AsyncAPI documentation , from AsyncAPI specifications, developers can automatically generate:\nScaffolding code and project structure Client libraries in multiple languages Server stubs and handlers Interactive documentation Per eBay\u0026rsquo;s article on AsyncAPI :\n\u0026ldquo;For industry adoption, however, it\u0026rsquo;s not just the richness of the specification that matters — it\u0026rsquo;s also the tooling that comes with it. As with OpenAPI, AsyncAPI tooling incorporates powerful visualizers that allow architects and engineers to collaborate on the design.\u0026rdquo;\nContract Testing: Per eBay :\n\u0026ldquo;Microcks is an open source Kubernetes mock-and-test framework that supports AsyncAPI.\u0026rdquo;\n6.3 Governance and Compliance API governance is the structured framework that manages and enforces policies for API deployment, security, observability, and lifecycle operations. Per Traefik Labs :\n\u0026ldquo;Open observability standards like OpenTelemetry guarantee compatibility with various backend API analytics platforms and prevent vendor lock-in. They provide a way to decouple your observability strategy from your gateway strategy.\u0026rdquo;\nData Contracts: Per Monte Carlo\u0026rsquo;s guide on data contracts :\n\u0026ldquo;A data contract is an agreement between a service provider and data consumers\u0026hellip; But here\u0026rsquo;s what makes it powerful: this agreement is implemented in code, not just documented in prose.\u0026rdquo;\nBest practices from Monte Carlo :\n\u0026ldquo;Place the contract under version control in Git, just like code. Store it in a schema registry for easy access.\u0026rdquo; \u0026ldquo;Embed checks in the CI/CD pipeline of the data producer. If an engineer tries to deploy a change that violates the contract, automated tests should fail.\u0026rdquo; \u0026ldquo;Add circuit breakers in data ingestion too. If incoming data doesn\u0026rsquo;t match the contract, stop it from flowing into the warehouse.\u0026rdquo; Runtime Enforcement: Per Traefik Labs :\n\u0026ldquo;API gateways provide control points for runtime API governance, enabling governance teams to enforce common runtime standards such as authentication and authorization, granular API access, rate limiting and throttling, input validation, and monitoring.\u0026rdquo;\n7. Recommendations for Product Development 7.1 Core Capabilities to Build Capability Why It Matters Priority OpenAPI Support Industry standard for REST APIs. Enables design-first development, auto-documentation, and code generation. (OpenAPI Initiative ) High AsyncAPI Support Essential for event-driven APIs. MuleSoft, Boomi, SAP, and others are adopting it. Protocol-agnostic design. (AsyncAPI ) High CloudEvents Format CNCF-graduated standard. Used by Microsoft, SAP, AWS for billions of messages daily. Ensures event interoperability. (CNCF ) Medium-High Schema Registry Centralized schema management with compatibility checking. Prevents breaking changes and enables schema evolution. (Confluent ) High OTel Semantic Conventions Emerging standard for observability data naming. Enables consistent monitoring across polyglot environments. MuleSoft and Boomi now support. (OpenTelemetry ) Medium EDI Support Required for B2B in retail, healthcare, logistics. X12 and EDIFACT remain foundational for supply chain. Medium (industry-dependent) Schema Validation Runtime validation prevents bad data from corrupting systems. Essential for API governance and security. High 7.2 Key Questions for Product Strategy When evaluating schema support in integration products, consider:\nDiscoverability\nCan developers find and understand available schemas through a catalog or portal? Is there a searchable registry of events, channels, and message types? Evolution\nDoes the platform support schema versioning with compatibility checking? Can schemas evolve without breaking existing consumers? Validation\nCan schemas be validated at design time AND runtime? Are validation errors surfaced clearly to developers? Interoperability\nDoes the platform support industry standards or lock users into proprietary formats? Can data flow to/from other platforms without custom translation? Tooling\nCan developers generate code, documentation, and tests from schema definitions? Is there IDE integration for schema authoring? 7.3 Implementation Roadmap Phase 1: Foundation\nImplement OpenAPI support for REST APIs Add schema validation at the gateway layer Create a basic schema/API catalog Phase 2: Event-Driven\nAdd AsyncAPI support for event-driven APIs Implement CloudEvents format for event interoperability Integrate with or build a schema registry Phase 3: Advanced Governance\nAdd schema evolution with compatibility checking Implement contract testing capabilities Build discoverability features (search, catalog, portal) Phase 4: Enterprise Scale\nAdd EDI support for B2B scenarios Implement multi-format support (Avro, Protobuf, JSON Schema) Enable federation across multiple registries/gateways 8. Conclusion The enterprise integration landscape is converging on open standards for schemas and message types. Organizations like CNCF (CloudEvents), the AsyncAPI Initiative, the OpenAPI Initiative, and the OpenTelemetry project are driving standardization that benefits the entire ecosystem.\nFor product managers building integration platforms, the message is clear: products that understand, respect, and utilize schema definitions will deliver superior value through interoperability, governance, and evolvability. This includes not just the data contracts themselves (OpenAPI, AsyncAPI, CloudEvents, Schema Registry) but also how those integrations are monitored (OpenTelemetry Semantic Conventions). Those that ignore these standards will find themselves creating integration chaos rather than solving it.\n\u0026ldquo;Change is inevitable. Breakage doesn\u0026rsquo;t have to be.\u0026rdquo;\n— Estuary on Schema Evolution (source )\nThe question is not whether to support these standards, but how quickly and comprehensively to do so.\nReferences and Further Reading Specifications Standard URL OpenAPI https://www.openapis.org/ AsyncAPI https://www.asyncapi.com/ CloudEvents https://cloudevents.io/ CloudEvents GitHub Spec https://github.com/cloudevents/spec Confluent Schema Registry https://docs.confluent.io/platform/current/schema-registry/ JSON Schema https://json-schema.org/ OpenTelemetry Semantic Conventions https://opentelemetry.io/docs/specs/semconv/ OpenTelemetry Messaging Conventions https://opentelemetry.io/docs/specs/semconv/messaging/ OpenTelemetry CloudEvents Conventions https://opentelemetry.io/docs/specs/semconv/cloudevents/ Platform Documentation Platform Resource URL MuleSoft AsyncAPI Event-Driven API Release Notes https://docs.mulesoft.com/release-notes/platform/event-driven-api MuleSoft Implementing AsyncAPI Specifications https://docs.mulesoft.com/anypoint-code-builder/imp-asyncapi MuleSoft AsyncAPI Specifications in Design Center https://docs.mulesoft.com/design-center/design-async-api MuleSoft OpenTelemetry Support https://docs.mulesoft.com/mule-runtime/latest/otel-support Boomi API Management https://boomi.com/platform/api-management/ Boomi Universal API Management Blog https://boomi.com/blog/how-universal-api-management-simplifies-data-modeling/ Boomi Platform Overview https://boomi.com/platform/ WSO2 API Manager https://wso2.com/api-manager/ WSO2 API Manager Architecture https://apim.docs.wso2.com/en/latest/get-started/apim-architecture/ Solace Event Portal https://solace.com/products/portal/ Solace Schema Registry Best Practices https://docs.solace.com/Schema-Registry/schema-registry-best-practices.htm Confluent Schema Evolution https://docs.confluent.io/platform/current/schema-registry/fundamentals/schema-evolution.html Industry Perspectives Source Topic URL Capital One AsyncAPI in Event-Driven Architecture https://www.capitalone.com/tech/software-engineering/asyncapi-event-driven-architecture/ eBay AsyncAPI 2.0 Adoption https://innovation.ebayinc.com/stories/asyncapi-2-0-enabling-the-event-driven-world/ Solace MuleSoft Event-Driven Integration https://solace.com/blog/mulesoft-embracing-event-driven-integration/ CNCF CloudEvents Graduation Announcement https://www.cncf.io/announcements/2024/01/25/cloud-native-computing-foundation-announces-the-graduation-of-cloudevents/ The New Stack CloudEvents Deep Dive https://thenewstack.io/cncf-cloudevents-a-lil-message-envelope-that-travels-far/ MuleSoft Evented APIs Blog https://blogs.mulesoft.com/api-integration/strategy/evented-apis-for-event-driven-challenges/ Monte Carlo Data Contracts Explained https://www.montecarlodata.com/blog-data-contracts-explained/ Decodable Schema Evolution in CDC Pipelines https://www.decodable.co/blog/schema-evolution-in-change-data-capture-pipelines Estuary Real-Time Schema Evolution https://estuary.dev/blog/real-time-schema-evolution/ Airbyte Mastering Schema Evolution https://airbyte.com/data-engineering-resources/master-schema-evolution OpenTelemetry ECS and OTel SemConv Convergence https://opentelemetry.io/blog/2023/ecs-otel-semconv-convergence/ OpenTelemetry HTTP Conventions Declared Stable https://opentelemetry.io/blog/2023/http-conventions-declared-stable/ Elastic ECS Contribution to OpenTelemetry FAQ https://www.elastic.co/blog/ecs-elastic-common-schema-otel-opentelemetry-faq Grafana Database Observability with OTel Semantic Conventions https://grafana.com/blog/2025/06/06/database-observability-how-opentelemetry-semantic-conventions-improve-consistency-across-signals/ EDI Resources Source Topic URL SEEBURGER ANSI X12 Overview https://www.seeburger.com/resources/good-to-know/what-is-ansi-x12 Zenbridge EDIFACT vs X12 https://zenbridge.io/insights/edifact-vs-x12/ EDI2XML EDI X12 Explained https://www.edi2xml.com/blog/edi-x12-explained-codes-standards-and-real-world-impact/ GraceBlood Understanding EDI Standards https://graceblood.com/blog/understanding-edi-standards-x12-vs-edifact/ Medium Introduction to B2B EDI https://medium.com/another-integration-blog/introduction-to-b2b-edi-x12-091bb4d4fe7e Case Studies Organization Use Case Source TransferGo AsyncAPI for event-driven microservices AsyncAPI Case Studies Slack Data streaming API documentation AsyncAPI Case Studies IBM Company-wide event catalog AsyncAPI Case Studies Port of Rotterdam Infrastructure monitoring with AsyncAPI AsyncAPI Case Studies Document generated from industry research including sources from MuleSoft, Boomi, WSO2, Solace, Confluent, CNCF, Capital One, eBay, and other industry leaders. All quotes and statistics are attributed with source links.\nCFP Schemas Are the Interface: Why Structure Matters More in the Age of AI Agents Schemas have always been the right way to define how systems exchange data. OpenAPI for REST, AsyncAPI for events, Protobuf for cross-language contracts, CloudEvents for message envelopes. The arguments for contract-first development are well established: type safety, code generation, breaking change detection, independent service evolution. But adoption has always been optional. Teams could skip the schema, ship raw JSON, and deal with the consequences later.\nAI agents changed the math. When an LLM decides which tool to invoke, it reads a schema. MCP tool definitions are JSON Schema. OpenAI function calling requires typed parameter declarations. Google\u0026rsquo;s Agent Development Kit generates callable tools directly from OpenAPI specifications, using the summary and description fields to guide tool selection. The model doesn\u0026rsquo;t read your source code or your documentation site. It reads the schema. A missing description is no longer a documentation gap. It is a capability gap. An imprecise type is no longer a developer annoyance. It is a source of hallucination.\nThis talk connects the decades-old case for schema-driven development to the new reality of AI agent tooling. You\u0026rsquo;ll see how MCP, OpenAI Structured Outputs, and Google ADK all depend on the same foundation (machine-readable, semantically annotated schemas), why Google is pushing gRPC and Protobuf as a transport layer for MCP, and what this means for how you design APIs and define contracts going forward. The schema is no longer just the contract between your services. It is the interface your AI agents reason against.\nTalk outline Slated at ~25 min runtime but can be adapted to a lightning talk or deep dive\nSchemas have always mattered (3 min) Quick tour of the landscape: OpenAPI, AsyncAPI, Protobuf, CloudEvents, Schema Registry. The traditional case for contract-first: type safety, code generation, compatibility, independent evolution. Why adoption remained optional for most teams despite the clear benefits. Enter the agents (5 min) How LLM function calling works: the model reads a schema, selects a tool, emits structured JSON. The orchestrator validates and executes. MCP\u0026rsquo;s tool definition format: name, description, inputSchema. The schema is the executable contract. Demo: a simple MCP server with two tools. Show the JSON Schema the model sees. Walk through how the model picks the right tool based on description quality. The new correctness requirements (7 min) Description quality drives tool selection accuracy. Vague descriptions cause wrong tool picks. Demo: same tool with a good vs. bad description, showing the model\u0026rsquo;s behavior change. Type enforcement prevents injection. Protobuf\u0026rsquo;s strict typing and JSON Schema\u0026rsquo;s constrained outputs both validate at the serialization layer, catching malformed inputs before execution. Schema drift breaks automated pipelines. In agent workflows, schema changes cascade. Practitioners identify schema drift as the top cause of broken automations. Missing fields are capability gaps. If your OpenAPI spec has no operationId or summary, Google ADK generates a tool the agent cannot reliably use. The ecosystem convergence (5 min) Google contributing gRPC transport to MCP (February 2026). Protobuf\u0026rsquo;s strict typing + MCP\u0026rsquo;s semantic descriptions = both layers matter. OpenAI Structured Outputs: constraining model generation at the decoding layer against a JSON Schema, not just prompting for JSON. Google ADK\u0026rsquo;s OpenAPIToolset: one OpenAPI spec in, callable tools out. The spec quality determines agent quality. Emerging standards: OASF (Open Agentic Schema Framework), agents.json, LangChain Agent Protocol. Different organizations arriving at the same conclusion. Practical adoption (5 min) Start with what you have: improve your OpenAPI descriptions and add missing operationId fields. This is the lowest-effort, highest-impact change for AI readiness. Add Protobuf for type safety across service boundaries. Use Buf for linting, breaking change detection, and code generation. Schema registries for versioning. The same compatibility strategies (backward, forward, full) that protect human consumers now protect agent consumers. The contract-first workflow you\u0026rsquo;d use for APIs now serves both humans and agents. The tooling (Buf, OpenAPI generators, schema registries) is already built. Target audience Backend developers, API designers, platform engineers, and anyone building or integrating with AI agent tooling (MCP servers, function calling, agentic workflows). Also relevant for engineering leaders evaluating how their API strategy intersects with AI adoption.\nTakeaways AI agents read your schemas at runtime to decide what to do. Schema quality directly determines agent accuracy. MCP, OpenAI function calling, and Google ADK all depend on machine-readable, semantically annotated schemas. The model is now a consumer of your API contract. Contract-first development (Protobuf, OpenAPI, JSON Schema) is no longer just good practice for teams. It is a prerequisite for reliable AI-powered automation. The tooling ecosystem built for human developers (Buf, schema registries, OpenAPI generators) already serves agents too. The investment in schemas pays off twice. References MCP Specification Google Pushes for gRPC Support in Model Context Protocol (InfoQ, Feb 2026) Google Cloud: gRPC as a Native Transport for MCP OpenAI Structured Outputs OpenAI Function Calling Google ADK: OpenAPI Tools Simon Willison on LLM Schemas (Feb 2025) Open Agentic Schema Framework (OASF) Comparing 7 AI Agent-to-API Standards (Nordic APIs) Building Agents with OpenAPI: LangChain vs Haystack (Speakeasy) ","date":"February 18, 2026","permalink":"/posts/schema-standards-learning-guide/","summary":"Exploring OpenAPI, AsyncAPI, CloudEvents, Schema Registry, OpenTelemetry Semantic Conventions, EDI, and More - A Learning Resource for Technical Product Managers","title":"Understanding Schemas and Message Types in Enterprise Data Integration","type":"posts"},{"content":" AI-assisted: I developed the arguments and references in this post. Claude helped organize and expand the draft.\nIt isn\u0026rsquo;t about he protocol it is about the schema. You don\u0026rsquo;t need to give up REST to get the benefits of Protobufs.\nContract-first is a different mindset Some developers naturally think in terms of defining the contract before writing any code. The service, its methods, its message types, all defined up front. Then tooling generates everything from that definition.\nSOAP did this with XML/WSDL, gRPC does it with Protobuf. Different eras, same instinct. The wire format is not the concern. The contract is the source of truth. Define the contract, generate the code, trust the process. JSON is the default JSON is human-readable, universally supported, and works everywhere. Curl, browser dev tools, Postman, every language. That\u0026rsquo;s a real strength and it\u0026rsquo;s why JSON became the default format for APIs. Nothing wrong with that.\nJSON doesn\u0026rsquo;t have a built-in schema. The shape of the data is implied by the code. Schema definitions (OpenAPI, Swagger, JSON Schema) are often added as an afterthought. Most teams build code-first and generate docs later. Contract-first (or \u0026ldquo;design-first\u0026rdquo;) requires defining the schema before writing business logic. That\u0026rsquo;s a different workflow that takes looking at the problem different. Most developers have only used one approach JSON\u0026rsquo;s ubiquity means fewer developers actively think about their serialization or schema choice. It\u0026rsquo;s just what you use.\nBaaS platforms (Firebase, Supabase) abstract the API layer. The data format is hidden behind an SDK call. Heavy frameworks and vibe coding generate the API layer. AI defaults to JSON because that\u0026rsquo;s what the training data contains. JSON is the default because so many before already chose it. The well worn path, not an active decision. Most developers haven\u0026rsquo;t had the opportunity to try contract-first development. Bringing attention to Protobufs Rising tides float all boats. The goal is to teach developers another way to define their APIs that opens up more opportunities.\nA .proto file defines message types and service contracts in one place. Single definition, everything generated from it. The schema is the source of truth, not the code. You define the contract first, then generate clients, servers, and docs from it. Any object that leaves the application boundary (DTOs, API responses, events, messages) can be defined as a Protobuf message from the start. That schema becomes shareable across any service that needs it, regardless of language. Two Go services, a Python worker, and a TypeScript frontend can all generate types from the same .proto file. Tools like buf (bufbuild/buf) handle linting, formatting, and breaking change detection against that definition. The schema has its own development lifecycle. protovalidate (bufbuild/protovalidate) adds validation rules directly in the .proto file. The contract enforces its own constraints. Protobuf forward and backward compatibility reduces breaking changes, lets teams upgrade independently. The transport protocol is a separate decision. gRPC, REST (via gRPC-gateway or Envoy transcoding), ConnectRPC (connectrpc/connect-es) all work from the same .proto definition. JSON still flows through REST gateways, mapped from the proto contract. More OSS projects should offer Protobuf-defined APIs alongside their existing interfaces. Normalizes contract-first thinking. Larger projects with explicit service contracts get easier to maintain over time, not harder Buf\u0026rsquo;s argument: the real reason to use Protobuf isn\u0026rsquo;t performance, it\u0026rsquo;s that schema-driven development reduces integration failures Learning to trust the contract is the real shift DevRel strategy Start with open source tools that already have developer-facing APIs. Contribute Protobuf definitions to real projects.\nJust creating a PR brings awareness. Maintainers and watchers see it, the conversation starts. Don\u0026rsquo;t create PRs with AI-generated code. Authenticity earns credibility with maintainers. OpenTelemetry (OTel) is a strong starting point. gRPC is already a core transport, data model defined in Protobuf. Metrics, traces, and logs are exactly the kind of structured data contracts where Protobuf excels Buf\u0026rsquo;s tooling-first DevRel approach (CLI, BSR, Buf Slack) has done more for adoption than performance benchmarks Do it live Stream the contributions. Demystifies the Protobuf workflow for developers who have never touched a .proto file.\nQuestions answered in real time, friction points become visible and solvable in front of an audience VODs, clips, and writeups continue educating after the stream ends Each contribution becomes a case study for introducing schema-driven APIs to existing projects It is time to reconsider Protobuf - Blog and Talk The ROI on Protobuf has never been better. The tooling, the ecosystem, and the developer experience have all changed since the common criticisms were written.\nA direct rebuttal to the feedback collected in I Reviewed 1,000s of Opinions on gRPC . Those criticisms were valid when written. Today each one can either be reversed or is much more nuanced: \u0026ldquo;Tooling is immature.\u0026rdquo; buf CLI, BSR, VS Code and IntelliJ plugins, Postman gRPC support. The tooling gap has closed. \u0026ldquo;Build process overhead.\u0026rdquo; buf generate is one command. Remote plugins via BSR mean you don\u0026rsquo;t need protoc installed locally. \u0026ldquo;Debugging is hard.\u0026rdquo; buf curl, grpcurl, Postman, and ConnectRPC\u0026rsquo;s JSON mode all make inspection straightforward. \u0026ldquo;No browser support.\u0026rdquo; ConnectRPC works over standard HTTP without a proxy. This is solved. \u0026ldquo;Load balancing is tricky.\u0026rdquo; Service meshes (Istio, Linkerd, Envoy) handle gRPC-aware load balancing natively now. \u0026ldquo;It\u0026rsquo;s over-engineering for non-Google scale.\u0026rdquo; The argument was always about performance. The real argument is about schemas. You don\u0026rsquo;t need Google scale to benefit from clear contracts. The deciding factor should be: is this project going to have schemas? Schemas almost always matter. If data crosses a boundary, a schema makes that boundry well defined and easier to maintain. Protobuf is the most capable schema language available and can be used anywhere schemas are found. API contracts, event schemas, data pipelines, configuration, inter-service messages, and even REST APIs. CFP It Is Time to Reconsider Protobuf Protobuf adoption remains low despite years of maturity, but not for the reasons most developers think. The real barrier is not complexity or tooling; it is that most developers have only ever worked with JSON and never had a reason to choose something different. Protobuf does not ask you to compete with JSON on its home turf. It asks you to think about your interfaces differently.\nThe real case for Protobuf isn\u0026rsquo;t serialization speed. It\u0026rsquo;s contract-first development. One .proto file drives type generation across every language in your stack, schema drift becomes a lint error, and breaking changes get caught before they ship. Modern tooling (buf, ConnectRPC, protovalidate) has removed every historical friction point. This talk covers the practical path to adopting Protobuf without abandoning REST or JSON where they already work.\nThe tooling story has changed significantly. buf handles linting, formatting, and breaking change detection in CI. ConnectRPC works over plain HTTP without a proxy. protovalidate puts validation rules directly in the schema. Postman, VS Code, and IntelliJ all have native support.\nThis talk walks through the most common criticisms of Protobuf and gRPC, acknowledges where they came from, and shows what has changed. You\u0026rsquo;ll see the current state of the tooling, a practical workflow for adopting Protobuf schemas in an existing project, and why the real ROI is in schema-driven development, not raw throughput.\nTalk outline 15 min lightning. Budget ~13 min of content, 2 min of slack. No live demo, screen recordings only. No Q\u0026amp;A built in.\nThe hook: watch what happens around every new protocol (1 min) MCP is the newest protocol most of this room has touched. It\u0026rsquo;s JSON-RPC, and in December 2025 its maintainers decided transports should be pluggable rather than blessing new official ones. Google withdrew its dedicated gRPC transport proposal, because with pluggable transports it wasn\u0026rsquo;t needed. Buf then published a full Protobuf mapping of MCP anyway, the same way they\u0026rsquo;ve shown up next to every other protocol. That\u0026rsquo;s the pattern worth noticing. Schema tooling arrives next to whatever the new thing is, because the schema question is independent of the protocol question. Hold that thought. You remember the bumpy years (2 min) Most of this room formed an opinion about Protobuf and gRPC somewhere between 2016 and 2020, and it was rough: protoc in your build, no browser story, thin IDE support, \u0026ldquo;you\u0026rsquo;re not Google.\u0026rdquo; Source it honestly from I Reviewed 1,000s of Opinions on gRPC so nobody thinks you\u0026rsquo;re strawmanning. Those complaints were correct. Say plainly that you\u0026rsquo;re not here to argue the complaints were wrong. You\u0026rsquo;re here because the thing they were about has changed underneath them. Do not attempt a point-by-point rebuttal. It reads as defensive and burns the clock. The reframe: Protobuf is a schema language (4 min) This is the talk. Everything before it is setup, everything after is evidence. Any object crossing an application boundary is a candidate: DTOs, events, API responses, config. Kafka and Confluent Schema Registry treat Protobuf as a first-class schema with zero gRPC involved. Strongest proof that schema and transport are separate decisions, and most of the room hasn\u0026rsquo;t connected it. Forward and backward compatibility as a built-in property, not a versioning convention you maintain by hand. Land the line: the deciding question is not \u0026ldquo;do I need gRPC,\u0026rdquo; it\u0026rsquo;s \u0026ldquo;does data cross a boundary here.\u0026rdquo; It became first-class while you weren\u0026rsquo;t looking (4 min) The point of this section is accumulation, not any single item. Nobody announced \u0026ldquo;Protobuf is ready now.\u0026rdquo; It arrived one platform at a time, starting at the infrastructure edge and working inward toward the code you write. Run it as a timeline slide, roughly one line each, fast: 2020: AWS ALB routes gRPC natively with end-to-end HTTP/2. Confluent Schema Registry makes Protobuf first-class alongside Avro. 2022: Postman ships gRPC support. The \u0026ldquo;you can\u0026rsquo;t just poke at it\u0026rdquo; objection loses its tool of choice. 2023: Kubernetes 1.27 promotes native gRPC health probes to GA. No sidecar, no wrapper binary. 2025: gRPC Swift 2 lands as a full async/await rewrite. Tonic is donated into the gRPC project under CNCF and becomes the official Rust implementation. protovalidate reaches v1.0. 2026: Protobuf gets a real language server. Spring Boot 4.1 ships first-party gRPC auto-configuration for server, client, and test. Then slow down and show exactly one of them. The LSP is the right pick: a twenty-second screen recording of go-to-definition and completion inside a .proto does more than any claim you can make out loud. Second artifact if time allows: buf breaking failing in CI on a renamed field. One screenshot, no narration. The line that ties it together: none of these were Protobuf asking for special treatment. Each one was a platform deciding a schema-defined contract was worth supporting directly. Monday morning (1.5 min) Pick one DTO that already exists in your codebase. Define it as a .proto. Generate it alongside your current JSON. Change no transport, delete no code. Add buf lint and buf breaking to CI. That\u0026rsquo;s the whole first step. The point is that adoption is additive. Nobody has to approve a migration. Close (0.5 min) One slide: the reframe restated, a QR code to the reference list, done. Cut for time, deliberately Protobuf Editions. Correct, current, and a nuance trap. Invites \u0026ldquo;is proto3 dead\u0026rdquo; and costs three minutes. The four gRPC streaming types. Transport detail, undercuts the schema thesis. Benchmarks and payload-size numbers. Arguing performance concedes the frame. BSR and remote plugins. Real value, but it\u0026rsquo;s a second-step concern. The remaining four criticisms from the konfig list. Blog material, not stage material. Slide outline 18 slides, ~10 minutes narrated, 15 minute slot. Same beats as the talk outline. The difference is that no slide is a sentence I\u0026rsquo;m about to say out loud.\nThe rule: three slide types, nothing in between.\nType What\u0026rsquo;s on screen Word budget Artifact A real screenshot or recording of the thing itself A label. 3–6 words. Diagram Boxes, arrows, a timeline. Something with a shape. Node labels only. No sentences. Punctuation One short phrase, large, alone 5 words maximum If a slide needs a sentence to make sense, it\u0026rsquo;s a slide I haven\u0026rsquo;t figured out yet. The audience is either reading or listening, never both, and I\u0026rsquo;d rather they listen.\nThe spine of the deck is one visual callback. Slide 4 shows a .proto file open in a 2018 editor: gray monospace, no highlighting, no help. Slide 13 shows the same file under the language server. I don\u0026rsquo;t explain the pairing. Nine slides apart, it lands on its own.\nAct 1: The pattern (3 slides, ~1 min) 1. Title — 10s · punctuation Talk title, name, handle. Nothing else.\n2. Three real screenshots — 30s · artifact Side by side, in order: the MCP \u0026ldquo;pluggable transports\u0026rdquo; blog post header, the SEP-1352 GitHub issue with its Closed badge visible, and the bufbuild/mcp-proto repo header. Dates under each. On screen: the three dates only. Say: MCP decided transports should be pluggable rather than blessing new ones. That made Google\u0026rsquo;s dedicated gRPC proposal unnecessary, so Google closed it themselves. Buf mapped the protocol to Protobuf anyway.\n3. Every protocol grows a schema — 25s · diagram Five columns, protocol on top and its schema language underneath. SOAP → WSDL. REST → OpenAPI. gRPC → Protobuf. GraphQL → SDL. MCP → (blank, then fills in). On screen: ten words, all of them proper nouns. Say: This keeps happening. The protocol shows up first and a schema language grows next to it, because the schema question is separate from the protocol question. That\u0026rsquo;s the whole talk.\nAct 2: The bumpy years (3 slides, ~1.5 min) 4. A .proto file in 2018 — 30s · artifact Screenshot of a .proto open in an editor of the era. Flat gray monospace, no syntax highlighting, no go-to-definition, no error squiggles. Nothing but text. On screen: \u0026ldquo;2018\u0026rdquo;. Say: This is what the file looked like. Your editor had no idea what it was.\n5. What people actually said — 45s · artifact Screenshots of real HN and Reddit comments. Cropped tight, four of them, timestamps visible. Say: Read two aloud, let the others land silently. Why screenshots: typed-out quotes read as strawmen no matter how accurate they are. A screenshot with a timestamp can\u0026rsquo;t be accused of paraphrasing, and it solves the sourcing problem at the same time.\n6. All true — 15s · punctuation On screen: \u0026ldquo;All of it was true.\u0026rdquo; Say: I\u0026rsquo;m not here to tell you those were wrong. The thing they were about changed underneath them.\nAct 3: The reframe (5 slides, ~4 min) — this is the talk 7. Thesis — 20s · punctuation On screen: \u0026ldquo;Protobuf is a schema language.\u0026rdquo; Say: Nothing. Let it sit for two seconds.\n8. One file, many outputs — 50s · diagram A single .proto in the center. Arrows fanning out to Go structs, TypeScript types, a Python client, a Kafka topic, and generated docs. The center node is small; the fan is the point. On screen: six node labels. Say: One definition. Every consumer generates from it. Nobody hand-writes the shape twice, and nobody\u0026rsquo;s copy drifts.\n9. A Kafka topic with a schema — 40s · artifact Screenshot of a schema registry UI showing a topic with schema type PROTOBUF. No arrows, no annotation. On screen: \u0026ldquo;No gRPC in this picture.\u0026rdquo; Say: Protobuf as a first-class schema type since 2020, alongside Avro. Topics, evolution, breaking-change checks. This is the proof that schema and transport are separate decisions.\n10. Old reader, new message — 50s · diagram Two writers and one reader. v2 writer emits a message with an extra field; the v1 reader parses it fine, with the unknown field drawn grayed out and passing straight through. On screen: \u0026ldquo;v1\u0026rdquo;, \u0026ldquo;v2\u0026rdquo;, \u0026ldquo;ignored\u0026rdquo;. Say: Add a field, pick an unused number, ship it. Compare that to a version key in a JSON envelope and a convention nobody wrote down.\n11. The question — 40s · punctuation On screen: \u0026ldquo;Does data cross a boundary here?\u0026rdquo; Say: Say it, pause, move on. Don\u0026rsquo;t elaborate. It returns on the last slide.\nAct 4: It became first-class (4 slides, ~3 min) 12. Timeline — 60s · diagram Horizontal axis, 2020 to 2026. Markers with platform logos, revealed left to right. Underneath, an arrow labeled \u0026ldquo;inward\u0026rdquo; running from infrastructure to your framework. On screen: years, logos, five one-word themes. Say: Name each year and what tier it hit. AWS and Confluent, then Postman, then Kubernetes, then the language implementations, then Spring Boot and the IDEs. Don\u0026rsquo;t read the items — the shape is the argument.\n13. The same file, now — 40s · artifact Twenty second silent capture: go-to-definition across files, then completion inside a message. The same .proto from slide 4. On screen: \u0026ldquo;2026\u0026rdquo;. Say: Almost nothing. Let people notice the callback themselves.\n14. buf breaking fails the build — 30s · artifact Screenshot of the red CI check on a renamed field. On screen: nothing beyond the screenshot. Say: Schema drift is a build failure now, not an incident.\n15. No launch day — 20s · punctuation On screen: \u0026ldquo;Nobody announced this.\u0026rdquo; Say: Each platform independently decided a schema-defined contract was worth supporting. That\u0026rsquo;s what maturing actually looks like, and it\u0026rsquo;s easy to miss.\nAct 5: What to do (3 slides, ~1.5 min) 16. One new file — 60s · diagram A file tree, before and after. The after side has exactly one new .proto and a generated directory, marked green. Everything else is unchanged and gray. On screen: the tree, plus buf lint and buf breaking. Say: Pick a DTO you already have, define it as a proto, generate next to your existing JSON, add the two checks to CI. No transport changes. No deletions.\n17. Additive — 20s · punctuation On screen: \u0026ldquo;Nobody has to approve a migration.\u0026rdquo;\n18. Close — 20s · punctuation On screen: the question again, plus a QR to the reference list.\nAssets to produce Real work is here now, not in the slide text. Roughly in order of how long each takes.\nScreen recording, Protobuf LSP: go-to-definition then completion, 20s, silent, large font. Must use the same file as the 2018 screenshot. Period screenshot, a .proto with no editor support. Find a 2018-era editor build or disable the plugin and match the theme of the era. Four cropped comment screenshots from HN and Reddit, timestamps visible, sourced from the konfig post\u0026rsquo;s citations. Screenshot, buf breaking failing a CI check on a renamed field. Screenshot, schema registry UI showing a topic with schema type PROTOBUF. Three screenshots for slide 2: MCP transports post, closed SEP-1352 issue, mcp-proto repo. Diagram, one .proto fanning out to five consumers. Diagram, v2 message read by a v1 reader with the unknown field grayed. Diagram, protocol-to-schema-language columns. Diagram, horizontal timeline with logos and the \u0026ldquo;inward\u0026rdquo; arrow. Diagram, before/after file tree. QR code to the published reference list. Slides deliberately not made A logo grid of supporting languages. The timeline says the same thing with a shape and a direction. Any benchmark chart. A performance number on screen concedes the frame. A gRPC streaming diagram. Transport detail that pulls against the schema thesis. A Protobuf Editions slide. Invites \u0026ldquo;so is proto3 dead?\u0026rdquo; and costs three minutes. A gateway and proxy architecture diagram. Different, longer talk. Any slide containing a sentence I am about to say out loud. Target audience Polyglot developers, API designers, and platform engineers who evaluated Protobuf or gRPC in the past and decided against it, or who have only ever worked with JSON APIs.\nTakeaways The Protobuf tooling ecosystem has matured to the point where the old friction points simply don\u0026rsquo;t exist for most use cases Protobuf is a schema language, not just a serialization format or a gRPC dependency You can adopt Protobuf schemas without giving up REST or JSON Schema-driven development reduces integration failures regardless of project scale References Schema-driven development and the case for Protobuf The real reason to use Protobuf is not performance - Buf\u0026rsquo;s argument that schema-driven development (not speed) is why Protobuf matters API Design-First vs. Code First - Stoplight on why defining contracts before writing code reduces rework Protobuf compatibility and schema evolution Protobuf Language Guide: Updating A Message Type - Official Google docs on safe field changes and wire compatibility Protobuf Dos and Don\u0026rsquo;ts - Official best practices for evolving schemas Backward and Forward Compatibility with Protocol Buffers - Practical walkthrough of Protobuf\u0026rsquo;s extensibility model Bridging REST and gRPC grpc-gateway - Generates a REST reverse-proxy from .proto service definitions Envoy gRPC-JSON Transcoder - Infrastructure-level REST-to-gRPC translation gRPC-Web - Browser client support for gRPC via proxy Tooling and ecosystem Buf Schema Registry (BSR) - Hosted Protobuf registry with dependency management and generated SDKs Protobuf finally has LSP support. You\u0026rsquo;re welcome. - Buf ships the first production-grade Protobuf language server (Jan 2026), closing the \u0026ldquo;IDE support is bad\u0026rdquo; criticism Introducing the next generation of the Buf CLI - v2 config format, monorepos as first-class citizens, buf config migrate protovalidate - Validation rules declared in the schema; reached v1.0 in September 2025 OpenTelemetry Protocol (OTLP) Specification - Real-world example of Protobuf as an industry-standard wire format opentelemetry-proto - The actual .proto definitions for OTel\u0026rsquo;s data model Protobuf Editions Protobuf Editions Overview - Official docs on replacing syntax = \u0026quot;proto3\u0026quot; with edition = \u0026quot;2024\u0026quot; Protobuf Editions are here: don\u0026rsquo;t panic - Buf\u0026rsquo;s take: editions are a feature-flag refactor, most users should stay on proto3 for now Protobuf Editions explained - Practical walkthrough of what changes and what doesn\u0026rsquo;t Protobuf changes announced June 27, 2025 - Edition 2024 release timeline ConnectRPC and the browser story Making gRPC more approachable with ConnectRPC - Introduction to Connect\u0026rsquo;s HTTP/1.1 + JSON approach ConnectRPC: Where is it now? - Two-year retrospective (May 2026) covering remote plugins, LSP, protovalidate, FauxRPC, OpenAPI generation Why Smart Teams Are Betting on ConnectRPC Over Standard gRPC - Side-by-side comparison of ergonomics and debuggability Connect RPC vs. Google gRPC: Conformance Deep Dive - How Connect implementations measure against the gRPC spec First-class support, platform by platform The timeline behind the talk\u0026rsquo;s central claim: Protobuf and gRPC support arrived incrementally across the industry, starting at the infrastructure edge and working inward.\nALB support for end-to-end HTTP/2 and gRPC - AWS, October 2020. gRPC-aware routing, health checks, and access logs at the load balancer Protobuf Schema Serializer and Deserializer - Confluent Platform 5.5 (2020) made Protobuf a first-class schema type alongside Avro Postman Now Supports gRPC - January 2022 open beta, GA with Postman v10 Kubernetes 1.24: gRPC container probes in beta - Alpha in 1.23, beta in 1.24, GA in 1.27. Native gRPC health checking with no wrapper binary Introducing gRPC Swift 2 - February 2025. Full async/await rewrite with pluggable transports and client-side load balancing gRPC-Rust Preview Release - Tonic donated into the gRPC project under CNCF as the official Rust implementation Spring Boot 4.1 Adds gRPC Auto-Configuration - June 2026. Server, client, testing, SSL, security, and health indicators all auto-configured Supported languages - The current roster, 15+ languages Language and framework integration The Ultimate Guide to Spring gRPC - Deep dive on Spring Boot 4.1\u0026rsquo;s first-party gRPC support: all four RPC types, error mapping, interceptors, metadata, deadlines, TLS, testing Getting Started with Spring gRPC in Spring Boot 4.1 - Shorter intro to the auto-configuration and client injection story (June 2026) spring-grpc - The Spring project itself gRPC-Rust Roadmap - Where the Rust implementation is headed after the Tonic donation Protobuf beyond RPC: events and data pipelines Why a Protobuf schema registry? - Buf\u0026rsquo;s case for schemas as a governed artifact, not a build detail Bufstream schema providers - Broker-side schema awareness for Kafka topics, Confluent Schema Registry API compatible How to use Protobuf with Apache Kafka and Schema Registry - Hands-on walkthrough Running it in production Six Lessons from Production gRPC - Operational friction points and how teams work around them Running gRPC at Scale: Lessons From the Frontlines of Production - gRPConf India talk recap on multi-region throughput, load balancing, observability Developer sentiment and community discourse I Reviewed 1,000s of Opinions on gRPC - Synthesizes developer opinions from Reddit, HN, Twitter, and YouTube HN: Can somebody please explain why we would use gRPC? - Candid practitioner discussion on gRPC\u0026rsquo;s value and friction points HN: A detailed comparison of REST and gRPC - Real-world experience reports on the REST vs gRPC tradeoff Protobuf/gRPC momentum Google Pushes for gRPC Support in Model Context Protocol - Google Cloud contributing gRPC transport to Anthropic\u0026rsquo;s MCP (Feb 2026) A gRPC Transport for the Model Context Protocol - Google Cloud\u0026rsquo;s own announcement (Jan 2026), with the full argument for binary encoding, mTLS, and method-level authorization SEP-1352: Add gRPC as a transport - Withdrawn by its Google authors on Jan 14, 2026 in favor of pluggable transports. The discussion thread is the real value: a live argument over whether schema-first belongs in a JSON-RPC protocol The Future of MCP Transports - The decision that made SEP-1352 unnecessary. gRPC ships as a custom transport, not an official one The 2026 MCP Roadmap - \u0026ldquo;We are not adding more official transports this cycle.\u0026rdquo; Keeps the momentum claim honest bufbuild/mcp-proto - Buf\u0026rsquo;s Protobuf mapping of the MCP protocol, open-sourced in response to the SEP-1352 discussion. A worked example of retrofitting a schema onto a JSON-native protocol gRPC and AI: A Powerful Partnership - How LLMs shorten the proto-authoring and test-generation loop gRPConf 2025 - A dedicated conference is itself a signal of ecosystem health; gRPConf 2026 follows Sept 3 at the Computer History Museum ","date":"February 14, 2026","permalink":"/posts/proto-popularity/","summary":"Protobuf schemas get overlooked because most developers have only ever worked with JSON","title":"Why are Protobufs not more popular? - WIP","type":"posts"},{"content":" AI-assisted: I designed the architecture described here and built the proof of concept. Claude helped draft and structure the writeup. All benchmarks and code examples are from real implementation work.\nYour application already delivers business data to users through purpose-built endpoints, schemas, and query infrastructure. But both operators and users want to know more. What happened to my file? When did that job run? How has usage changed over the last quarter? The kind of questions that live outside your core domain model.\nThe traditional answer is to build more. More search endpoints, more storage schemas, more application nodes handling queries, and more database capacity to collect and serve results. Costs balloon. Complexity compounds. And you still end up with a narrow view of what the application is actually doing.\nIt doesn\u0026rsquo;t have to work that way. Your application is already capturing this information. It\u0026rsquo;s sitting inside your OpenTelemetry traces. Every span carries business context: files processed, documents modified, messages delivered, errors encountered. The problem is that it\u0026rsquo;s buried under thousands of internal spans that mean nothing outside of DevOps, locked inside tooling that nobody wants to teach end users to navigate.\nflowchart LR subgraph INPUT[\"📡 Ingestion\"] collector[\"OTEL Collector100% of spans\"] end subgraph FILTER[\"🔍 Pipeline\"] pipeline{\"Business EventFilter\"} end subgraph SERVE[\"⚡ Query Tiers\"] direction TB warm[\"Warm PathDuckDB · 23ms0.7% of data\"] cold[\"Cold PathJSONL on S3 · 5s8% of data\"] archive[\"ArchiveFull OTEL · 60s100% retained\"] end noise([\"🗑️ Infrastructure Noise92% discarded\"]) collector ==\u003e|\"all tracesvia OTLP\"| pipeline pipeline ==\u003e warm pipeline ==\u003e cold pipeline -.-\u003e archive pipeline -.-\u003e noise style INPUT fill:#1a1a2e,stroke:#16213e,color:#eee style FILTER fill:#16213e,stroke:#0f3460,color:#eee style SERVE fill:#0f3460,stroke:#533483,color:#eee style noise fill:none,stroke:#666,color:#999,stroke-dasharray: 5 5 style warm fill:#064e3b,stroke:#10b981,color:#ecfdf5 style cold fill:#78350f,stroke:#f59e0b,color:#fef3c7 style archive fill:#7f1d1d,stroke:#ef4444,color:#fef2f2 style collector fill:#1e3a5f,stroke:#3b82f6,color:#dbeafe style pipeline fill:#312e81,stroke:#818cf8,color:#e0e7ff This post walks through building a user-facing audit log archive that filters business events from raw OTEL telemetry, stores them in object storage, and makes them searchable through a type-safe API. It\u0026rsquo;s accessible, safe, and cheap. No application overhead, no per-query costs, no vendor lock-in. Just industry-standard tooling (OpenTelemetry, Protobuf, DuckDB) wired together in a way that serves the people who actually use your software.\nThe Problem: Observability Data vs. User Data Application observability and user-facing audit logs are usually treated as separate concerns. Observability tools like Datadog and Splunk capture everything (HTTP requests, database queries, cache hits, internal retries) and surface it to DevOps teams. Audit logs are typically a separate system built for compliance or support teams, not the application users themselves: a database table, a dedicated logging service, or a third-party compliance tool.\nBut if your application is already instrumented with OpenTelemetry, the data from both flows through the same pipeline. Every business transaction that matters to your users (a file upload, a document modification, a payment processed) already exists as OTEL spans with rich attributes. The challenge is filtering out the observability noise and capturing only the data relevant to end users.\nThis creates two problems:\nVolume and cost. A platform processing thousands of transactions per day generates gigabytes of raw OTEL data monthly. Most of it is internal infrastructure spans that no user will ever see. Storing and querying all of it through managed platforms (BigQuery at $6.25/TB scanned, Datadog\u0026rsquo;s retention pricing) gets expensive fast, and most of that cost goes toward data users don\u0026rsquo;t need.\nTwo query patterns. Users browsing their recent activity need fast, filtered, paginated results, ideally under 100 milliseconds. But when they need to drill into a specific transaction for a support ticket or compliance audit, they need the full event timeline with all attributes and context, even if it takes a few seconds. No single storage tier serves both needs economically.\nThe approach described here solves both problems: filter OTEL telemetry to business-relevant events, store them in tiered storage matched to query patterns, and expose them through a search API that users can query directly.\nWhy DuckDB? DuckDB is an open-source, in-process SQL database designed for analytical queries. It\u0026rsquo;s often described as \u0026ldquo;SQLite for analytics\u0026rdquo; because it runs embedded in your application with no external server to manage. What makes it useful for audit log archives specifically is its ability to read structured files directly from object storage.\nWith the httpfs extension, DuckDB can query JSONL and Parquet files stored in S3, GCS, or any S3-compatible store, such as MinIO. It handles on-the-fly gzip decompression, reads Hive-partitioned directory layouts, and uses columnar execution with vectorized processing (the same techniques used by purpose-built analytics engines).\nIn a proof of concept against real production OTEL data, the performance gap was significant:\nData Tier Query Time Monthly Size Reduction vs Raw Full OTEL JSONL (archive) 52s 3.8 GB — Business-filtered JSONL 29s 2.1 GB 45% Audit-event JSONL 5s 400 MB 92% DuckDB warm path 23ms 26 MB 99.3% The first row is everything the OTEL Collector captures: every internal span, every HTTP request, every database query. The third row is just the business events that matter to users. Filtering out observability noise reduces data volume by 92%.\nThe warm path (a DuckDB file containing denormalized audit summaries) is 2,200x faster than scanning raw OTEL JSONL. That\u0026rsquo;s the difference between a sub-second API response and a minute-long wait.\nDuckDB is also free. No per-query cost, no scan-based billing, no retention tiers. It runs wherever your application runs: in a container, a serverless function, or on a developer\u0026rsquo;s laptop. The only cost is the compute your application is already using.\nArchitecture: Hot / Warm / Cold Instead of forcing all queries through a single storage layer, the architecture uses tiered storage matched to query patterns:\nflowchart TD collector[OTEL Collector] pipeline[Data Pipeline] duckdb[(DuckDB File)] index[(Index — MongoDB)] storage[(Object Storage — GCS/S3/MinIO)] api[Audit Log Search API] consumers[Consumers — UI, Support Tools, APIs] collector -- all traces via OTLP --\u003e pipeline pipeline -- warm path --\u003e duckdb pipeline -- trace lookup --\u003e index pipeline -- audit events JSONL --\u003e storage duckdb --\u003e api index --\u003e api storage --\u003e api api --\u003e consumers The important piece is the data pipeline sitting between the OTEL Collector and storage. It receives the full firehose of observability data and filters it down to business-relevant events, the audit trail that users and support teams actually need. Internal infrastructure spans (HTTP middleware, database connection pools, cache operations) get discarded or routed to a separate observability store.\nHot: In-memory cache or server-sent events. Sub-10ms responses for repeated queries and real-time browser updates. Standard application-level caching.\nWarm: A DuckDB file containing denormalized audit summaries: transaction IDs, statuses, timestamps, durations, and event counts. Written by the ingestion processor, read concurrently by the search API. DuckDB supports concurrent readers with a single writer, so there are no conflicts. Query times are under 100ms, fast enough for paginated user-facing UIs.\nCold: Filtered JSONL files on object storage, organized by date using Hive-style partitioning (year=YYYY/month=MM/day=DD/). These contain the full event detail for each business transaction. DuckDB reads them directly via httpfs, handling compression and schema detection on its own. Response times range from 2 to 10 seconds, which is acceptable for drill-down views and support investigations.\nArchive: The full, unfiltered OTEL JSONL export from the Collector. The compliance-grade record of everything, including internal spans. Query times range from 10 to 60+ seconds and are used only for deep debugging and regulatory audits.\nA trace index (in MongoDB, or any fast key-value store) provides O(1) lookups to locate which archive file contains a specific transaction, avoiding full scans when drilling into a known trace ID.\nThe same query engine (DuckDB) serves all tiers. The warm path queries a local file. The cold path queries remote JSONL. The API code is nearly identical; only the file path changes.\nTechnology Choices The technology decisions here were driven by accessibility, adaptability, and cost. An audit log system needs to scale with user demand, not just infrastructure load, so the choices lean toward tools that are easy to adopt and cheap to run.\nDuckDB is open-source, embeddable, and requires zero infrastructure. It reads JSONL and Parquet from any S3-compatible storage. Anyone who knows SQL can write queries against it. There\u0026rsquo;s no cluster to manage, no license to negotiate, and no vendor lock-in. If DuckDB disappeared tomorrow, the JSONL files in object storage would still be just files, queryable by any tool that reads JSON.\nOpenTelemetry Collector is the data collection layer. It\u0026rsquo;s vendor-neutral by design: it receives traces via OTLP (gRPC or HTTP), applies filtering and routing, and exports to any destination. The Collector\u0026rsquo;s file exporter writes JSONL with configurable rotation and compression. Using OTEL means the audit log pipeline isn\u0026rsquo;t a separate instrumentation system; it piggybacks on the observability instrumentation your application already has.\nNode.js was chosen for the search API for developer familiarity. More developers on a typical team can contribute to and maintain a Node.js/TypeScript service than one written in Java, Go, or Rust. It deploys to containers, serverless functions, and VMs with minimal friction. If the cost or scaling needs change, the service can be ported to Cloud Functions or Cloud Run with minimal rework.\nProtobuf provides the schema contract across the pipeline. The same .proto definitions generate types for every language in the system (Java for the ingestion processor, TypeScript for the search API, and any other consumer). The choice here is about correctness, not performance. For an audit log system where data integrity matters, a single source of truth for domain types eliminates an entire category of bugs.\nAll four are open-source, widely adopted, and replaceable. The architecture doesn\u0026rsquo;t depend on any single vendor\u0026rsquo;s proprietary technology.\nThe Protobuf \u0026amp; gRPC Data Pipeline OTEL Is Built on Protobuf OpenTelemetry is more than a set of SDKs. It\u0026rsquo;s a data format specification built on Protocol Buffers. The OTLP protocol defines trace, metric, and log schemas as Protobuf messages. When the OTEL Collector receives a span via gRPC, it arrives as a structured ExportTraceServiceRequest proto message with strongly typed fields for trace IDs, span kinds, attributes, events, and status codes.\nProtobuf here acts as a schema contract. Field types, field numbers, and field names are guaranteed not to change after a specification reaches stability. The OTEL proto stability guarantees ensure backward compatibility: new fields can be added, but existing fields are never removed or renamed. For an audit log archive that may need to query data written years ago, stability matters.\nSchema Synchronization Across Applications When an ingestion processor (written in Java) and a search API (written in TypeScript) both need to understand what an \u0026ldquo;audit event\u0026rdquo; looks like, you have two choices:\nManual mapping. Each service defines its own types and manually maps between them. Every schema change requires coordinated updates across services, and drift is inevitable. Generated types from a shared proto. Define the schema once in .proto files, generate language-specific types, and share the definitions across services. The second approach eliminates drift by construction. Tools like Buf make this workflow practical by handling linting, breaking change detection, and code generation across languages from a single buf.gen.yaml configuration. Instead of wiring together protoc plugins manually, Buf manages the full generation pipeline.\nConsider an audit domain with 40+ event types spanning file processing, document modification, messaging, request/response cycles, error handling, scheduling, and retry logic:\n// Defined once in the event.proto enum EventType { EVENT_TYPE_UNSPECIFIED = 0; // File events (1-9) EVENT_TYPE_FILE_RECEIVED = 1; EVENT_TYPE_FILE_PROCESSED = 2; EVENT_TYPE_FILE_SENT = 4; // Document events (10-19) EVENT_TYPE_DOCUMENT_RECEIVED = 10; EVENT_TYPE_DOCUMENT_MODIFIED = 13; // Message events (20-29) EVENT_TYPE_MESSAGE_RECEIVED = 20; EVENT_TYPE_MESSAGE_PROCESSED = 22; // ... 40+ event types total } Running Buf\u0026rsquo;s code generation with protoc-gen-es (for TypeScript) and protoc-gen-java (for Java) configured in buf.gen.yaml produces type-safe enums in both languages from this single file. Buf\u0026rsquo;s breaking change detection (buf breaking) also catches any accidental modifications to stable field numbers or enum values before they reach a build. Adding a new auditable event type means adding one line to the proto. Both services pick it up on their next build.\nDeriving OTEL Event Names from Proto Enums One practical benefit of proto-first design: OTEL span event names can be derived algorithmically from enum names instead of maintained in a separate mapping table:\n// Enum name: FILE_RECEIVED // → lowercase: file_received // → prepend prefix: myapp.event.file_received // → OTEL span event name export function toOtelEventName(type: EventType): string { const enumName = EventType[type]; return \u0026#34;myapp.event.\u0026#34; + enumName.toLowerCase(); } // Reverse: parse OTEL event name back to proto enum export function fromOtelEventName(otelEventName: string): EventType { const enumName = otelEventName.substring(\u0026#34;myapp.event.\u0026#34;.length).toUpperCase() as keyof typeof EventType; return EventType[enumName] ?? EventType.UNSPECIFIED; } So:\nThe proto enum is the single source of truth for event classification. No separate mapping files to maintain. Event name derivation is testable and deterministic. End-event classification (which events mark a transaction as complete) is defined as a Set\u0026lt;EventType\u0026gt; alongside the enum, not scattered across services. This is also how the data pipeline decides which spans are audit-relevant: if a span contains events matching known EventType names, it\u0026rsquo;s a business event. Everything else is infrastructure noise.\nHow Protobuf Connects to DuckDB The proto-defined schema determines three things in the search layer:\nWhat gets extracted from raw OTEL JSON? The domain mapper knows which OTEL attributes to pull because the proto defines the audit event fields. How it maps to DuckDB. Proto field names map to DuckDB column aliases in query CTEs. What\u0026rsquo;s searchable in the API? The SearchFields configuration maps user-friendly names to OTEL attribute keys, which are themselves derived from proto-defined semantic conventions. A schema change in the proto propagates through the entire stack: the ingestion processor writes the new field, the search service extracts it, and the API exposes it to users. One proto update.\nImplementation Deep-Dive Data Layout: Hive-Partitioned JSONL on Object Storage The OTEL Collector\u0026rsquo;s file exporter writes JSONL files to object storage. Organizing them with Hive-style partitioning enables DuckDB to scan only the relevant date partitions:\ns3://audit-data/ year=2025/ month=01/ day=15/ events-001.json.gz events-002.json.gz day=16/ events-001.json.gz month=02/ ... DuckDB reads compressed JSONL directly, with no decompression step or ETL pipeline. A query for January 15th reads only the files in year=2025/month=01/day=15/.\nA practical caveat: DuckDB\u0026rsquo;s glob-based Hive partition filtering can be slower than expected. When given a broad glob pattern, DuckDB enumerates all directories first, then filters them, rather than walking only the matching directories. In benchmarks, this showed as ~23 seconds with a glob versus ~3 seconds with explicit paths. The fix is to build explicit per-day paths in application code instead of relying on a single broad glob:\n// Instead of: s3://bucket/year=*/month=*/day=*/*.json.gz // Build explicit paths for the requested date range: private buildGlobPathsForRange(startDate: string, endDate: string): string[] { const paths: string[] = []; for (let d = new Date(start); d \u0026lt;= end; d.setDate(d.getDate() + 1)) { const year = d.getFullYear(); const month = String(d.getMonth() + 1).padStart(2, \u0026#39;0\u0026#39;); const day = String(d.getDate()).padStart(2, \u0026#39;0\u0026#39;); paths.push( `s3://${bucket}/year=${year}/month=${month}/day=${day}/*.json*` ); } return paths; } This limits the scan to exactly the requested days, with a configurable maximum (e.g., 90 days) to prevent runaway queries.\nDuckDB Client: CLI Subprocess Pattern There are two ways to use DuckDB from Node.js: via native bindings or via a CLI subprocess. The CLI approach (spawning duckdb -json -c \u0026lt;sql\u0026gt;) has the advantage of greater stability.\nNote on Node.js bindings: The original duckdb npm package is deprecated as of DuckDB v1.4.x (Fall 2025). The recommended replacement is @duckdb/node-api, a high-level wrapper around @duckdb/node-bindings. If you\u0026rsquo;re starting a new project, evaluate the new bindings. The CLI subprocess pattern described here was chosen before the new bindings were available and remains a valid alternative for maximum compatibility.\nThe subprocess client configures S3 access via DuckDB\u0026rsquo;s httpfs extension, then runs the query:\nasync query\u0026lt;T\u0026gt;(sql: string): Promise\u0026lt;T[]\u0026gt; { const fullSql = ` INSTALL httpfs; LOAD httpfs; SET s3_endpoint=\u0026#39;${this.s3Endpoint}\u0026#39;; SET s3_access_key_id=\u0026#39;${this.accessKey}\u0026#39;; SET s3_secret_access_key=\u0026#39;${this.secretKey}\u0026#39;; SET s3_use_ssl=${this.useSsl}; SET s3_url_style=\u0026#39;path\u0026#39;; ${sql} `; return new Promise((resolve, reject) =\u0026gt; { const duckdb = spawn(\u0026#39;duckdb\u0026#39;, [\u0026#39;-json\u0026#39;, \u0026#39;-c\u0026#39;, fullSql]); let stdout = \u0026#39;\u0026#39;; duckdb.stdout.on(\u0026#39;data\u0026#39;, (data) =\u0026gt; { stdout += data.toString(); }); duckdb.on(\u0026#39;close\u0026#39;, (code) =\u0026gt; { if (code !== 0) return reject(new Error(`DuckDB failed`)); resolve(stdout.trim() ? JSON.parse(stdout) : []); }); }); } The -json flag tells DuckDB to emit results as a JSON array, which matches the API response format. Each query is a separate subprocess with no connection pooling or state leaks. If DuckDB crashes, it doesn\u0026rsquo;t take down the API process.\nThe s3_url_style='path' setting is important for S3-compatible stores like MinIO that use path-style URLs rather than virtual-hosted-style.\nUnnesting OTEL\u0026rsquo;s Nested Structure OTEL JSONL is deeply nested. A single line contains a resourceSpans array, each with scopeSpans, each with spans. Getting to individual spans requires three levels of unnesting.\nDuckDB\u0026rsquo;s unnest() function handles this well in CTEs (Common Table Expressions):\nWITH raw_data AS ( SELECT * FROM read_json( \u0026#39;s3://bucket/year=2025/month=01/day=15/*.json*\u0026#39;, format=\u0026#39;newline_delimited\u0026#39;, ignore_errors=true ) ), -- Level 1: unnest resourceSpans unnested AS ( SELECT unnest(resourceSpans) as rs FROM raw_data ), -- Level 2: unnest scopeSpans scope_spans AS ( SELECT unnest(rs.scopeSpans) as ss FROM unnested ), -- Level 3: unnest spans spans AS ( SELECT unnest(ss.spans) as span FROM scope_spans ) SELECT span.traceId as trace_id, span.spanId as span_id, span.name as span_name, CAST(span.startTimeUnixNano AS BIGINT) as start_nanos, CAST(span.endTimeUnixNano AS BIGINT) as end_nanos, span.attributes as attrs FROM spans The ignore_errors=true flag matters in production. Without it, a single malformed JSON line causes the entire query to fail.\nExtracting Attributes with list_filter and list_extract OTEL attributes are stored as an array of key-value pairs, not a flat map:\n{ \u0026#34;attributes\u0026#34;: [ { \u0026#34;key\u0026#34;: \u0026#34;http.method\u0026#34;, \u0026#34;value\u0026#34;: { \u0026#34;stringValue\u0026#34;: \u0026#34;POST\u0026#34; } }, { \u0026#34;key\u0026#34;: \u0026#34;http.status_code\u0026#34;, \u0026#34;value\u0026#34;: { \u0026#34;intValue\u0026#34;: \u0026#34;200\u0026#34; } }, { \u0026#34;key\u0026#34;: \u0026#34;myapp.transaction.id\u0026#34;, \u0026#34;value\u0026#34;: { \u0026#34;stringValue\u0026#34;: \u0026#34;abc-123\u0026#34; } } ] } Extracting a specific attribute requires filtering the array by key, then accessing the value. DuckDB\u0026rsquo;s list_filter() and list_extract() functions handle this:\n-- Extract a specific attribute value from the OTEL attributes array (list_extract( list_filter(attrs, x -\u0026gt; x.key = \u0026#39;myapp.transaction.id\u0026#39;), 1 )).value.stringValue as transaction_id This reads as: \u0026ldquo;filter the attributes array to entries where key equals \u0026lsquo;myapp.transaction.id\u0026rsquo;, take the first match, and extract its stringValue.\u0026rdquo;\nFor integer attributes (like HTTP status codes), swap stringValue for intValue. For booleans, boolValue. The value type matches the OTEL attribute type, which is defined by the proto schema.\nDynamic Search with Configurable Fields Instead of hardcoding which OTEL attributes are searchable, the search API uses a field configuration that maps user-friendly names to internal OTEL attribute keys:\ninterface SearchFieldDefinition { name: string; // API field name: \u0026#39;httpStatus\u0026#39; label: string; // UI label: \u0026#39;HTTP Status\u0026#39; otelKey: string; // OTEL key: \u0026#39;http.response.status_code\u0026#39; valueType: \u0026#34;string\u0026#34; | \u0026#34;int\u0026#34; | \u0026#34;bool\u0026#34; | \u0026#34;double\u0026#34;; category: \u0026#34;http\u0026#34; | \u0026#34;database\u0026#34; | \u0026#34;messaging\u0026#34; | \u0026#34;audit\u0026#34; | \u0026#34;service\u0026#34; | \u0026#34;error\u0026#34;; description: string; // Tooltip text } const SEARCH_FIELDS: SearchFieldDefinition[] = [ { name: \u0026#34;url\u0026#34;, label: \u0026#34;URL\u0026#34;, otelKey: ATTR_URL_FULL, // from @opentelemetry/semantic-conventions valueType: \u0026#34;string\u0026#34;, category: \u0026#34;http\u0026#34;, description: \u0026#34;Full request URL\u0026#34;, }, { name: \u0026#34;httpStatus\u0026#34;, label: \u0026#34;HTTP Status\u0026#34;, otelKey: ATTR_HTTP_RESPONSE_STATUS_CODE, valueType: \u0026#34;int\u0026#34;, category: \u0026#34;http\u0026#34;, description: \u0026#34;HTTP response status code\u0026#34;, }, // ... 18 fields across HTTP, Database, Messaging, Audit, File, Service, Error ]; OTEL attribute keys are sourced from @opentelemetry/semantic-conventions where available, keeping them consistent with the broader OTEL ecosystem. Custom domain attributes use a namespaced convention (e.g., myapp.transaction.id).\nThe API exposes a /fields endpoint that returns all searchable fields with metadata:\n{ \u0026#34;fields\u0026#34;: [ { \u0026#34;name\u0026#34;: \u0026#34;httpStatus\u0026#34;, \u0026#34;label\u0026#34;: \u0026#34;HTTP Status\u0026#34;, \u0026#34;valueType\u0026#34;: \u0026#34;int\u0026#34;, \u0026#34;category\u0026#34;: \u0026#34;http\u0026#34;, \u0026#34;description\u0026#34;: \u0026#34;HTTP response status code\u0026#34; } ], \u0026#34;count\u0026#34;: 18 } A UI consuming this endpoint can dynamically generate filter controls (dropdowns, text inputs, numeric ranges) without hardcoding field lists. Adding a new searchable field means adding one entry to the configuration; the API, query builder, and field discovery all pick it up. This works well for user-facing audit log UIs where different audiences need different filters: an end user reviews their own recent activity, a support agent searches by customer ID, a compliance team filters by date range and document type, and all three work against the same API.\nAdvanced Search: Structured Filters + Free-Text The advanced search endpoint combines structured attribute filters with a free-text fallback. Each filter specifies a field, an operator, and a value:\n{ \u0026#34;startDate\u0026#34;: \u0026#34;2025-01-01\u0026#34;, \u0026#34;endDate\u0026#34;: \u0026#34;2025-01-31\u0026#34;, \u0026#34;filters\u0026#34;: [ { \u0026#34;field\u0026#34;: \u0026#34;httpStatus\u0026#34;, \u0026#34;operator\u0026#34;: \u0026#34;gte\u0026#34;, \u0026#34;value\u0026#34;: \u0026#34;400\u0026#34; }, { \u0026#34;field\u0026#34;: \u0026#34;serviceName\u0026#34;, \u0026#34;operator\u0026#34;: \u0026#34;eq\u0026#34;, \u0026#34;value\u0026#34;: \u0026#34;order-api\u0026#34; } ], \u0026#34;freeText\u0026#34;: \u0026#34;timeout\u0026#34; } The query builder translates each filter into a type-aware SQL condition. For string fields, it uses quoted comparison. For numeric fields, it validates and casts:\n// Type-aware SQL generation based on field definition switch (filter.operator) { case \u0026#34;eq\u0026#34;: return fieldDef.valueType === \u0026#34;string\u0026#34; ? `AND ${column} = \u0026#39;${escapedValue}\u0026#39;` : `AND ${column} = ${validateNumeric(filter.value)}`; case \u0026#34;contains\u0026#34;: return `AND CAST(${column} AS VARCHAR) ILIKE \u0026#39;%${escapedValue}%\u0026#39;`; case \u0026#34;gte\u0026#34;: return `AND ${column} \u0026gt;= ${validateNumeric(filter.value)}`; // ... gt, lt, lte, neq } The free-text fallback uses ILIKE across the full JSON-serialized attributes:\nWHERE (structured_conditions) OR attrs_json ILIKE \u0026#39;%timeout%\u0026#39; This gives users a \u0026ldquo;Google-like\u0026rdquo; search experience (type anything and find matching audit events) while also supporting precise, structured queries when they know exactly what they\u0026rsquo;re looking for.\nHandling Nanosecond Timestamps OTEL timestamps are in nanoseconds (as strings in JSON, since JavaScript\u0026rsquo;s Number type can\u0026rsquo;t safely represent them). DuckDB casts them to BIGINT, and the search service uses JavaScript\u0026rsquo;s BigInt for arithmetic:\nconst startNanos = BigInt(row.start_nanos); const endNanos = BigInt(row.end_nanos); const durationMs = Number((endNanos - startNanos) / BigInt(1_000_000)); const startTime = new Date(Number(startNanos / BigInt(1_000_000))).toISOString(); Nanosecond precision is preserved through the DuckDB query and converted to millisecond ISO timestamps for the API response.\nResults \u0026amp; Viability The proof of concept confirms DuckDB works as a query engine for user-facing audit log archives. Results against real production data:\nQuery Pattern Warm (DuckDB file) Cold (JSONL on GCS) Single transaction by ID 22ms 4-11s Event timeline for transaction 30ms N/A Filtered list with pagination 28ms N/A Date range (1 week) 24ms 15-38s Monthly aggregation 23ms 57-144s The warm path handles all interactive UI queries under 30ms, fast enough that the audit log feels instant. The cold path handles drill-down into full transaction detail in single-digit seconds for most cases. If data hasn\u0026rsquo;t been processed into the warm tier yet, the cold path still works, just slower.\nWho This Serves Since the architecture filters business events from OTEL observability data and exposes them through a searchable API, it serves both internal teams and end users:\nEnd users browsing their transaction history, checking the status of a file upload, or reviewing a document processing timeline, directly in the application UI Support teams investigating a customer\u0026rsquo;s issue by searching for their transaction ID or correlation ID across months of audit data Compliance teams auditing transaction records for specific document types, date ranges, or error conditions DevOps teams that can fall back to the full OTEL archive when they need infrastructure-level spans for debugging The field discovery endpoint lets new consumers build their own query UIs without coordinating with the API team. A customer-facing audit view, an internal support tool, and a compliance dashboard can all use the same API with different field filters.\nThe Cost Equation The entire stack runs on open-source software: DuckDB, OpenTelemetry, Node.js, and Protocol Buffers. On Kubernetes, the search API runs as a standard deployment with horizontal pod autoscaling. No per-query billing, no scan-based pricing, no minimum spend.\nThe data lives in object storage, the cheapest durable storage available from any cloud provider. A year of audit log data that might cost thousands in a managed observability platform costs a few dollars per month in GCS or S3 standard storage. And because the pipeline filters out infrastructure noise before archiving, the volume is 92% smaller than raw OTEL exports.\nFor the warm path, a DuckDB file containing months of denormalized audit summaries stays well under 100 MB, fitting in memory on the smallest container.\nWhat\u0026rsquo;s Next A few optimizations we\u0026rsquo;re looking at next:\nDuckDB\u0026rsquo;s read-through caching (announced in January 2026 ) can cache HTTP range requests locally, reducing repeated cold-path query times. Parquet conversion for older archives. DuckDB\u0026rsquo;s columnar format is faster than JSONL for analytical queries, and DuckDB can perform the conversion itself. Redis caching for frequently-accessed query patterns, moving from in-process cache to shared cache across API replicas. Clay Smith\u0026rsquo;s \u0026ldquo;Cheap OpenTelemetry Lakehouses\u0026rdquo; work covers a similar Parquet + DuckDB + Iceberg pattern for OTEL storage and arrived at many of the same conclusions independently. The convergence is a good sign: embedded SQL engines over open file formats on object storage seem like a solid pattern for anyone who needs to make telemetry data accessible to users, not just DevOps.\nCFP Your App Already Has an Audit Log. It\u0026rsquo;s Hiding in Your OpenTelemetry Data. Your application already captures every business transaction that matters to users. File uploads, document modifications, payment processing, status changes. It\u0026rsquo;s all there, inside your OpenTelemetry traces, buried under thousands of infrastructure spans that nobody outside DevOps will ever need.\nThe traditional answer is to build a separate audit log system: new storage schemas, new endpoints, new database capacity. The cost compounds and you still end up with a narrow view of what the application is actually doing. This talk presents a different approach. Filter business events from the OTEL telemetry you already collect, store them in tiered object storage, and make them searchable with DuckDB. No vendor lock-in, no per-query cost, no separate instrumentation. Just industry-standard tooling (OpenTelemetry, Protobuf, DuckDB) wired together in a way that serves the people who actually use your software.\nYou\u0026rsquo;ll see real proof-of-concept benchmarks: filtering out 92% of observability noise to isolate business events, a warm path that returns paginated results in 23ms, and a cold path that drills into full transaction detail from object storage in single-digit seconds. The entire stack runs on open-source software with data stored in plain JSONL files on S3/GCS, the cheapest durable storage available.\nTalk outline Slated at ~25 min runtime but can be adapted to a lightning talk or deep dive\nThe problem: observability data vs. user data (4 min) Your OTEL traces already contain the audit trail. The challenge is separating business events from infrastructure noise. Why building a separate audit log system is expensive and redundant when you\u0026rsquo;re already instrumented. Architecture: hot / warm / cold (6 min) Tiered storage matched to query patterns: DuckDB warm path for sub-100ms interactive queries, JSONL on object storage for drill-down, full OTEL archive for compliance. The data pipeline that filters, routes, and classifies spans using Protobuf-defined event types. DuckDB as the query engine (6 min) Why DuckDB: embedded, zero-infrastructure, reads JSONL and Parquet from S3 directly via httpfs. Unnesting OTEL\u0026rsquo;s nested structure with CTEs. Extracting attributes with list_filter and list_extract. The Hive-partition glob performance caveat and the fix. Protobuf as the schema contract (4 min) Single .proto definition generates types for Java ingestion and TypeScript search API. Deriving OTEL event names from proto enums. One schema change propagates through the entire stack. Results and practical considerations (5 min) Benchmark walkthrough: 23ms warm vs. 52s raw OTEL scans. 99.3% data reduction. Cost equation: open-source stack, object storage pricing, no per-query billing. Who this serves: end users, support teams, compliance, DevOps fallback. Target audience Backend engineers, platform engineers, and DevOps practitioners who work with OpenTelemetry and are looking for ways to make telemetry data useful beyond dashboards. Also relevant for teams evaluating alternatives to managed observability platforms for audit and compliance use cases.\nTakeaways OpenTelemetry traces already contain business-relevant audit data. Filtering out infrastructure noise reduces data volume by 92%. DuckDB can query JSONL and Parquet files directly from object storage with no external server, no license, and no per-query cost. A Protobuf schema contract keeps event classification consistent across ingestion and search services in different languages. Tiered storage (DuckDB warm path + JSONL cold path + full OTEL archive) matches query patterns to cost and performance needs without forcing all queries through a single tier. ","date":"February 4, 2026","permalink":"/posts/building-otel-archive-search-with-duckdb/","summary":"How to filter business events from raw OpenTelemetry telemetry, store them in tiered object storage, and make them searchable with DuckDB — no vendor lock-in and no per-query cost.","title":"Building a User-Facing Audit Log Archive on OpenTelemetry and DuckDB | Rough Draft","type":"posts"},{"content":"My job as a leader is to help you be as successful or more successful in your career as I have been in mine. Your growth can happen in spite of a bad leader or because of a good one. I aim to be the latter.\nHere is how I boil it down. I want to be good enough at my own job and fluent enough in the jobs of the people who report to me that I can spend my time deliberately and treat your time as the more valuable resource. Seniority makes my time look more expensive on paper, so I flip that on purpose. I multiply my time by investing it in helping you get better at using yours.\nWhat you can expect from me I trust you to manage your own calendar (or status). I run 1:1s for you, not for me. I keep notes in a private doc shared with you. Anything that shows up in a performance review will already be in that doc. No surprises. I plan work as far ahead as I can. I use loose planning structures and add processes only when they add more value than the time they take. I encourage growth days. Take a full day or split it in half. Whatever fits how you learn. What I expect from you Ask for help when you need it. Nobody on my team works alone. Use your judgment on what to escalate. I don’t need to approve or know about everything. The test I use: if you think I would be surprised to hear about it later, I should hear about it now. When you need PTO, decide for yourself to take a day, and let the team know. For anything longer than three days, give me a heads up, ideally two weeks out on the calendar. Principles I lead by I focus on the team’s success over individual heroics. We define what success looks like together, and I don’t measure it with vanity metrics (numbers that look good in a report but don\u0026rsquo;t reflect real outcomes). I take new ideas seriously. Ideas should be fostered and evaluated then decided on. I treat endless debate as a failure mode in itself. We should only aim to be better than we were yesterday. I don’t let us compare ourselves to larger teams. I lead with trust. I only add process when trust is missing, or when the team has grown past what trust alone can hold. Other good references (WIP) Lara Hogan, Set future performance reviews up for success now — on running 1:1s and shared notes so reviews contain no surprises. Dan McKinley\u0026rsquo;s talk Egoless Engineering ","date":"February 1, 2026","permalink":"/posts/effective_management/","summary":"Your growth can happen in spite of a bad manager, or because of a good one.the latter.","title":"Effective management","type":"posts"},{"content":" Last updated: April 2026\nOverview This layout delivers a 10G fiber backbone across three managed switch locations, with 2.5G copper to every endpoint. The design principle: run fiber between floors or distant rooms so the inter-location links never become a bottleneck, then fan out 2.5G to devices at each location.\nWhy this topology:\n10G SFP+ fiber between switches — fiber is immune to interference, supports long runs without signal degradation, and the 10G headroom leaves room for aggregated device traffic to grow. 2.5G to endpoints — sufficient for NAS transfers, 4K streaming, and multi-gigabit workloads without requiring expensive 10G NICs everywhere. Chain topology (Main → Upstairs → Far End) — matches the physical layout of most houses; each switch location gets a full 10G uplink to the next. A star back to the main rack would require more fiber runs. Three managed switch zones — Main rack (servers + gateway), Upstairs (workstations + AP), Far end (third equipment zone, e.g. garage or basement). Each zone has spare ports for additional equipment. UCG-Fiber as the management boundary — all routing, firewall, and VLAN policy lives at the gateway; switches carry tagged traffic but don\u0026rsquo;t make policy decisions. Equipment Device Location Budget option UniFi equivalent Modem / ONT Main rack ISP-provided — Gateway Main rack UniFi Cloud Gateway Fiber (UCG-Fiber) — Main switch Main rack SODOLA 10-port 2.5G (8× 2.5G + 2× 10G SFP+) ~$75 USW-Pro-Max-8-PoE (8× 2.5G + 2× 10G SFP+) ~$299 Upstairs switch Upstairs SODOLA 10-port 2.5G (8× 2.5G + 2× 10G SFP+) ~$75 USW-Pro-Max-8-PoE (8× 2.5G + 2× 10G SFP+) ~$299 Far-end switch Third zone SODOLA 8-port 2.5G (8× 2.5G + 1× 10G SFP+) ~$60 USW-Flex-2.5G (8× 2.5G + 1× 10G SFP+) ~$199 AP #1 Main floor Ubiquiti U7 Lite (WiFi 7, 2.5GbE, PoE 802.3af) — AP #2 Upstairs Ubiquiti U7 Lite (WiFi 7, 2.5GbE, PoE 802.3af) — AP #3 Third zone Ubiquiti U7 Lite (WiFi 7, 2.5GbE, PoE 802.3af) — SODOLA vs. UniFi switches: SODOLA switches are third-party managed switches — full 802.1Q VLAN and LACP support, but no UniFi controller integration. UniFi switches add controller-managed config, topology visibility, and native PoE (eliminating separate PoE injectors for the APs) at a significantly higher price. The gateway and APs are UniFi either way; the switch choice is where the cost/ecosystem tradeoff lives.\nTopology graph TD ISP[\"🌐 ISP Fiber — 3 Gbps\"] MODEM[\"📦 Modem / ONT — 10GE RJ45 out\"] UCG[\"🔒 UCG-Fiber — WAN: 10G RJ45 — LAN: 10G SFP+\"] SW1[\"🔀 Main Switch — SODOLA 10-port\\n8× 2.5G + 2× 10G SFP+\\nMain Rack\"] SW2[\"🔀 Upstairs Switch — SODOLA 10-port\\n8× 2.5G + 2× 10G SFP+\\nUpstairs\"] SW3[\"🔀 Far-End Switch — SODOLA 8-port\\n8× 2.5G + 1× 10G SFP+\\nThird Zone\"] AP1[\"📶 U7 Lite AP #1 — 2.5G + PoE injector\"] AP2[\"📶 U7 Lite AP #2 — 2.5G + PoE injector\"] AP3[\"📶 U7 Lite AP #3 — 2.5G + PoE injector\"] S1[\"🖥️ Server #1 — 2.5G\"] S2[\"🖥️ Server #2 — 2.5G\"] S3[\"🖥️ Server #3 — 2.5G\"] WS1[\"🖥️ Workstation #1 — 2.5G\"] WS2[\"🖥️ Workstation #2 — 2.5G\"] ISP --\u003e|fiber| MODEM MODEM --\u003e|\"10G copper Cat6A — in-rack\"| UCG UCG --\u003e|\"10G SFP+ DAC — in-rack\"| SW1 SW1 --\u003e|\"2.5G RJ45 — same floor\"| AP1 SW1 --\u003e|\"2.5G RJ45 — same floor\"| S1 SW1 --\u003e|\"2.5G RJ45 — same floor\"| S2 SW1 --\u003e|\"2.5G RJ45 — same floor\"| S3 SW1 --\u003e|\"10G SFP+ fiber — inter-floor run\"| SW2 SW2 --\u003e|\"2.5G RJ45 — same floor\"| WS1 SW2 --\u003e|\"2.5G RJ45 — same floor\"| WS2 SW2 --\u003e|\"2.5G RJ45 — same floor\"| AP2 SW2 --\u003e|\"10G SFP+ fiber — far-end run\"| SW3 SW3 --\u003e|\"2.5G RJ45 — same zone\"| AP3 VLAN Design All inter-switch fiber links carry 802.1Q trunks with all VLANs tagged. The UCG-Fiber enforces inter-VLAN routing and firewall policy at the gateway — switches forward tagged frames but apply no policy of their own.\nVLAN Name Purpose Port mode 10 Trusted Servers, workstations Access on server/workstation ports; tagged on all trunk links 20 IoT Smart home, cameras Access on IoT device ports; tagged on all trunk links 30 Guest Guest WiFi SSID Tagged on AP ports (SSID-to-VLAN mapping on U7 Lite); tagged on trunk links 99 Management Switch OOB management IPs Tagged on trunk links; no untagged access ports Trunk ports: All SFP+ inter-switch links (UCG→SW1, SW1→SW2, SW2→SW3) carry VLANs 10, 20, 30, 99 tagged.\nAP ports: Configured as trunk with VLAN 10 untagged (management) and VLANs 20/30 tagged. Each SSID on the U7 Lite maps to a VLAN — Trusted SSID → VLAN 10, IoT SSID → VLAN 20, Guest SSID → VLAN 30.\nFirewall policy at UCG-Fiber:\nIoT (VLAN 20) → blocked from Trusted (VLAN 10); internet allowed Guest (VLAN 30) → blocked from all internal VLANs; internet allowed Trusted (VLAN 10) → full access Port Mapping graph LR subgraph MODEM[\"Modem / ONT\"] M_OUT[\"10GE RJ45 OUT\"] end subgraph UCG[\"UCG-Fiber\"] UCG_WAN[\"WAN: 10G RJ45\"] UCG_LAN[\"LAN: 10G SFP+\"] end subgraph SW1[\"Main Switch — SODOLA 10-port\"] SW1_SFP1[\"SFP+ Port 1 (10G) — UCG uplink\"] SW1_SFP2[\"SFP+ Port 2 (10G) — inter-floor run\"] SW1_P1[\"RJ45 Port 1 (2.5G) — Server #1\"] SW1_P2[\"RJ45 Port 2 (2.5G) — Server #2\"] SW1_P3[\"RJ45 Port 3 (2.5G) — Server #3\"] SW1_P4[\"RJ45 Port 4 (2.5G) — AP #1\"] SW1_P5[\"RJ45 Port 5 (2.5G) — spare\"] SW1_P6[\"RJ45 Port 6 (2.5G) — spare\"] SW1_P7[\"RJ45 Port 7 (2.5G) — spare\"] SW1_P8[\"RJ45 Port 8 (2.5G) — spare\"] end subgraph SW2[\"Upstairs Switch — SODOLA 10-port\"] SW2_SFP1[\"SFP+ Port 1 (10G) — fiber from main\"] SW2_SFP2[\"SFP+ Port 2 (10G) — far-end run\"] SW2_P1[\"RJ45 Port 1 (2.5G) — Workstation #1\"] SW2_P2[\"RJ45 Port 2 (2.5G) — Workstation #2\"] SW2_P3[\"RJ45 Port 3 (2.5G) — AP #2\"] SW2_P4[\"RJ45 Port 4 (2.5G) — spare\"] SW2_P5[\"RJ45 Port 5 (2.5G) — spare\"] SW2_P6[\"RJ45 Port 6 (2.5G) — spare\"] SW2_P7[\"RJ45 Port 7 (2.5G) — spare\"] SW2_P8[\"RJ45 Port 8 (2.5G) — spare\"] end subgraph SW3[\"Far-End Switch — SODOLA 8-port\"] SW3_SFP1[\"SFP+ Port 1 (10G) — fiber from upstairs\"] SW3_P1[\"RJ45 Port 1 (2.5G) — AP #3\"] SW3_P2[\"RJ45 Port 2 (2.5G) — spare\"] SW3_P3[\"RJ45 Port 3 (2.5G) — spare\"] SW3_P4[\"RJ45 Port 4 (2.5G) — spare\"] SW3_P5[\"RJ45 Port 5 (2.5G) — spare\"] SW3_P6[\"RJ45 Port 6 (2.5G) — spare\"] SW3_P7[\"RJ45 Port 7 (2.5G) — spare\"] SW3_P8[\"RJ45 Port 8 (2.5G) — spare\"] end M_OUT ---|\"Cat6A patch — in-rack\"| UCG_WAN UCG_LAN ---|\"10G SFP+ DAC 0.5–1m — in-rack\"| SW1_SFP1 SW1_SFP2 ---|\"10G fiber — inter-floor run (measure distance)\"| SW2_SFP1 SW2_SFP2 ---|\"10G fiber — far-end run (measure distance)\"| SW3_SFP1 SW1_P1 ---|\"Cat6 patch — same floor\"| S1 SW1_P2 ---|\"Cat6 patch — same floor\"| S2 SW1_P3 ---|\"Cat6 patch — same floor\"| S3 SW1_P4 ---|\"Cat6 run to AP — same floor\"| AP1 SW2_P1 ---|\"Cat6 patch — same floor\"| WS1 SW2_P2 ---|\"Cat6 patch — same floor\"| WS2 SW2_P3 ---|\"Cat6 run to AP — same floor\"| AP2 SW3_P1 ---|\"Cat6 run to AP — same zone\"| AP3 Cable Runs Run From To Type Speed Location 1 Modem UCG-Fiber WAN Cat6A patch 10G In-rack 2 UCG-Fiber SFP+ Main Switch SFP+ 1 10G SFP+ DAC (0.5–1m) 10G In-rack 3 Main Switch SFP+ 2 Upstairs Switch SFP+ 1 OM3/OM4 fiber + 2× SFP+ MM modules 10G Inter-floor run 4 Main Switch RJ45 1 Server #1 Cat6 patch 2.5G Same floor 5 Main Switch RJ45 2 Server #2 Cat6 patch 2.5G Same floor 6 Main Switch RJ45 3 Server #3 Cat6 patch 2.5G Same floor 7 Main Switch RJ45 4 PoE Injector → AP #1 Cat6 solid copper 2.5G Same floor 8 Upstairs Switch RJ45 1 Workstation #1 Cat6 patch 2.5G Same floor 9 Upstairs Switch RJ45 2 Workstation #2 Cat6 patch 2.5G Same floor 10 Upstairs Switch RJ45 3 PoE Injector → AP #2 Cat6 solid copper 2.5G Same floor 11 Upstairs Switch SFP+ 2 Far-End Switch SFP+ 1 OM3/OM4 fiber + 2× SFP+ MM modules 10G Far-end run 12 Far-End Switch RJ45 1 PoE Injector → AP #3 Cat6 solid copper 2.5G Same zone SFP+ \u0026amp; Fiber Qty Item Where Used 1 10G SFP+ DAC 0.5–1m UCG-Fiber ↔ Main Switch (in-rack) 2 10G SFP+ multimode modules Main Switch SFP+ 2 + Upstairs Switch SFP+ 1 (inter-floor run ends) 1 OM3/OM4 LC-LC duplex fiber (measure length) Main → Upstairs wall run 2 10G SFP+ multimode modules Upstairs Switch SFP+ 2 + Far-End Switch SFP+ 1 (far-end run ends) 1 OM3/OM4 LC-LC duplex fiber (measure length) Upstairs → Far-End wall run ","date":"January 1, 2026","permalink":"/posts/fiber_home_network/","summary":"10G backed home setup","title":"Home Network Layout","type":"posts"},{"content":" TL;DR Highlights I wrote the blog and demo for the MQTT to Kafka feature announcement . My Video Walkthrough for the new AsyncAPI feature support announcement . The team had a tight deadline, meaning I had to record this over a weekend and couldn\u0026rsquo;t re-recorded it to add more polish. I completely revitalized the current Zilla documentation site . It started as a Gitbook site (wayback link ). I improved the content and reworked all of the demos and examples. I am proud of the MQTT broker product feature guide . I managed multiple product examples with a single startup script and E2E tests that run in a GitHub Action . Blogs MQTT to Kafka feature announcement and Demo Community Updates and Blog Posts My personal blogs Videos Video Walkthrough for an AsyncAPI feature announcement Technical content video with Twitch Studio tutorial Documentation How I organize documentation . The Zilla docs I inherited when I started vs when I left . A product feature guide that I am proud of. Demos/Examples How I organize examples . I managed multiple product examples with single startup script and E2E tests that run in a GitHub Action . Created Product Demos on GitHub that can easily be run locally with a Compose file. Vendor specific options can be set from the options passed into the startup script. I restarted and expanded the Ballerina Exercism.io Track . ","date":"February 1, 2025","permalink":"/posts/devrel_content_highlights/","summary":"Higlighting some of the public things I have built or written down.","title":"My DevRel Content Highlights","type":"posts"},{"content":"The role of Developer Relations (DevRel) is still finding where it fits best in an org. DevRel can make a large impact working in multiple departments but this makes goal alignment difficult. While DevRel historically emerged from marketing departments, its true home lies within Product. Many others have said this, this is why I think they are right.\nDeveloper Experience is Product-Centric A developer has to choose one of many products that technically fit their need. The Developer Experience (DevEx) isn\u0026rsquo;t just about slick APIs or polished SDKs. It\u0026rsquo;s the sum of every interaction a developer has with your product above and beyond core features. Onboarding flows, error messages, documentation, and even community support are all direct extensions of the product itself. When DevRel sits within Product, it ensures that the developer\u0026rsquo;s journey is treated as a core part of the product strategy, not an afterthought.\nProduct Design Shapes Developer Experience A well-designed product minimizes friction for developers. For example:\nA confusing API schema forces developers to waste time deciphering endpoints. Poorly designed SDKs lead to frustration and abandonment. Inconsistent versioning erodes trust. DevRel teams embedded in Product can advocate for developer-centric design decisions early in the development cycle. They act as the voice of the end-user, ensuring that the product is built with empathy for that developer struggling to get it.\nWhy Marketing KPIs Conflict with Developer Success DevRel was born in marketing because developers were seen as a \u0026ldquo;hard-to-reach\u0026rdquo; audience. Traditional marketing tactics like lead generation, conversion rates, and click-through metrics don\u0026rsquo;t directly translate to developer finding success. Developer reaction to marketing can range from apathy to disdain and depending on the measures taken can erode the developers trust that the product has their needs in mind.\nDevRel\u0026rsquo;s primary mission is to build trust and empower developers which is fundamentally misaligned with marketing\u0026rsquo;s focus on acquisition and growth. When DevRel reports to Marketing, it risks becoming a lead-gen tool rather than a champion for developers.\nDocumentation: Written by DevRel; Guided by Engineers In the race to keep up with product improvements and user demands, DevRel and the Technical Writers within the team are the drivers of good documentation with Product Engineers navigating. Engineers are experts at building products, not explaining them. It is hard for an engineer deep in technical understanding to look at their own product from the eyes of the end user. Not to mention the soft skills needed and the time it takes to write good docs. Documentation requires a different skill set: clarity, empathy, and an understanding of diverse user personas. Technical writers on a DevRel team should own documentation. When documentation lives under DevRel (which itself aligns with Product), it becomes a living resource that evolves with the product.\nDevRel\u0026rsquo;s Role: Guiding Developers to the Right Resources DevRel\u0026rsquo;s primary goal is to guide developers to success, entertaining, generating leads, etc are importent but secondary.\nCreating tutorials and demos that highlight key features. Directing developers to relevant documentation. Acting as a feedback loop between users and Product. When DevRel and Product share goals, incentives align. For example, a Product team focused on adoption will prioritize features that solve real problems, while DevRel ensures developers can easily find and understand those features.\nMarketing vs. DevRel: Entertainment vs. Information Marketing\u0026rsquo;s Role: Entertain, engage, and attract. Think viral tweets, swag, and event sponsorships. DevRel\u0026rsquo;s Role: Inform, educate, and retain. Think in-depth workshops, troubleshooting guides, and API deep dives. Marketing creates the first impression; DevRel ensures the second, third, and hundredth interactions are meaningful. Imagine Marketing as a matchmaker and DevRel as the wingman: Marketing gets developers in the door, but DevRel ensures they stay.\nThe Path Forward: Integrate DevRel with Product To maximize developer success:\nMove DevRel under Product and align incentives with developer outcomes. Empower technical writers as first-class stakeholders in product development. Treat documentation as a core product feature, not a support artifact. Celebrate the onboarding experience making it just has memorable as the first impression. When DevRel and Product work in lockstep, developers win and so does your product.\n","date":"January 29, 2025","permalink":"/posts/devrel_belongs_in_product/","summary":"Many others have said this, this is why I think they are right.","title":"Developer Relations Belongs in Product","type":"posts"},{"content":"Searching for cloud native jobs on different job boards is hard. I decided to scan the CNCF Landscape projects homepage URLs for any that might have a careers page. I wrote a script to help.\nhttps://github.com/vordimous/cncf-job-pages Here are the US Remote friendly jobs pages that I am following for anyone else who doesn\u0026rsquo;t want to sort through them all.\naltinity.com apollographql.com astronomer.io buoyant.io bytebase.com chainguard.dev chronosphere.io civo.com clickhouse.com codefresh.io coder.com controlplane cortex.io crunchydata.com decodable.co densify.com diagrid.io endorlabs.com emqx.com firehydrant.com fusionauth gitlab.com gitpod.io goteleport.com grafana.com gremlin.com greptime.com hashicorp.com highlight.io honeybadger.io honeycomb.io infisical.com influxdata.com kentik.com kestra.io konghq.com liquibase.com mariadb.com newrelic.com ngrok.com observeinc.com odigos.io omnistrate.com orkes.io ory.sh pachyderm.com percona.com permify.co planetscale.com platform.sh platform9.com posit.co progress.com pulumi.com redis.io redpanda.com rely.io robusta.dev secondfront.com sematext.com siderolabs.com signoz solo.io splunk.com stack.io steamhaus.co sumologic.com supabase.com superorbital.io sysdig.com temporal.io tenable.com timescale.com upbound.io vercel.com vultr.com wundergraph.com ","date":"December 1, 2024","permalink":"/posts/cncf_jobs/","summary":"Searching for cloud native jobs on different job boards is hard.","title":"Looking for a CNCF Job is hard","type":"posts"},{"content":"As a developer, I want an easy way to run different samples and use the code to configure my own usecases.\nAcceptance criteria Single entry point to run all samples with a script or cli Each sample shouldn\u0026rsquo;t have extra complexity that distracts from the demonstrated use Each sample has a simple setup and teardown script that handles everything for the dev and explains how the setup is achieved A sample should be testable via a test script Sample variations are self contained and adjusted by changing defined variables Quickstart All resources are found in the repo\nLow barrier of entry, fast success, 5min proof of function Used to confirm app can do what it says hosted where possible Cookbooks All resources are found in a cookbook folder or repo\nTeaching focused Organized by use case Exhaustive explanation Used to guide Variable backend vendors Includes Install instructions Examples All resources are found in an examples folder or repo\nFeature focused Organized by function Short copy \u0026amp; paste oriented Used when adding or troubleshoot features Static backend vendors Demos All resources are found in a demos repo\nShow and Tell focused End to End, Real world scenarios Quick setup and walkthrough Used to showcase multiple functions \u0026amp; use cases hosted when possible ","date":"June 8, 2024","permalink":"/posts/organizing_examples/","summary":"How to organizing your Quickstart, Cookbooks, Examples, and Demos.","title":"Organizing your Quickstart, Cookbooks, Examples, and Demos","type":"posts"},{"content":"Any video about your products is good for the brand. Your company needs to be able to promote content that meets quality standards. This guide is a best practice reference sheet, and any guidelines should be encouraged but not forced into every production.\nPrimary Components Test the Setup Before recording or presenting, ensuring everything will work before committing to the entire recording is good practice.\nRecord 1-5 minutes of your presentation for review. Test lighting considering the time of day can change. Look and listen for any audio and video issues. Try out different presentation elements to make sure they fit in the recording. Show your Face A friendly live camera can add personality to a video. However, ensure it does distract the viewer from the content.\nDo not use virtual backgrounds; record in a well-lit room with a simple backdrop. Use a white/light-colored background. Dress for business casual if you don\u0026rsquo;t know the audience. Scripts are always good but do not try to read a script while recording. It will sound stiff and as if you\u0026rsquo;re reading a book. Check that it isn\u0026rsquo;t blocking anything important. Record in a Sandbox When presenting or recording a video, your machine is the studio.\nClose any unnecessary apps and mute any potential notifications. Separate computer profiles or at least separate browser profiles will give you a nice sandbox environment to present. This can also prevent accidentally sharing sensitive information. Capture a specific portion of your screen or minimize toolbars. Browser tabs and bookmarks can be a security risk and are distracting. When using a command prompt or terminal screen, position it in the middle of your laptop and make the fonts bigger. This allows the editor to zoom in on the Command prompt, allowing viewers to see a better view of what you are typing. 20% Live Pre-recording the bulk of a presentation will significantly improve quality and consistency.\n80% of the time, a live presentation could be pre-recorded. For example, most digital conferences stream recorded presentations with live interaction from the host. The 20% live part can be an intro, the host replying to live comments, and/or a live Q\u0026amp;A at the end. This will result in more audience interaction and a better experience. No Background noise Ambient office noises can be much louder than the speaker realizes.\nFind a quiet place and record an audio test, including making noises around and outside the room. When typing, don\u0026rsquo;t press the keys too hard. Silence all device notifications and mute all electronics. Simple Movement Consumers can\u0026rsquo;t follow your demo as fast as you can move through it.\nIf you have any mouse movements, consider making as few as possible and do it slowly. If you perform some action, always try to explain why you did it or the impact it will have. Try to minimize swapping tabs or pages, as it can cause tearing in the video recording or stream. NO Text In any visual presentation, text on the screen is a detriment.\nDO NOT read slides. They should be a visual aid to the talking points. Don\u0026rsquo;t overpopulate slides with too much text. Simple outlines, at most, should be used. Don\u0026rsquo;t talk about URLs or how to find other content. Instead, mention resources and add them in the description or pinned comment. On Brand Mention the brand at the beginning and end. Add how a viewer can find more information or follow our progress. Ask for comments as we will follow up later General The entire video should not exceed 7-8 minutes (5min is ideal); if it does exceed, structure it so that you can break it into parts of a series. Avoid long gaps between narration when recording your video. Avoid using too many technical words and jargon (if using, explain). Follow proper code formatting guidelines and best practices. For example, do not write poorly formatted code first, letting the IDE auto-format it. Deliverables Raw recording and audio (if you have recorded it separately) It will be posted on YouTube or other social media channels as we see fit. We add the branding, and if there is anything you need to edit, please share them in detail so it saves time in editing. Screen Capturing Software Twitch Recommended Streaming tools Mac – Quicktime Player Linux – Kazam Ubuntu - Kazam Make sure the resolution is 1920x1080p or higher. The aspect ratio should be 16:9.\nAnatomy Welcome (~30s) Consumer attention fades quickly. The intro will be the most watched section. Briefly describe the video, DO NOT read the title An overview of the overall structure of the video Branding: Mention the company and the social channels consumers can find more information on the topic Introduce yourself. It is an excellent transition into the body of the content. Ex. “Hi, I\u0026rsquo;m X, so let\u0026rsquo;s get to it.” Problem(⅓ of the runtime) Use an outsider\u0026rsquo;s perspective Mention common reasons Talk about 1 or 2 problems if possible Solution(⅔ of the runtime) Solve each problem Clearly describe each step Demonstrate the solution if possible Include references Summary(1-2min) Repeat both problem and solution in a brief summary Repeat Branding and ask for feedback Reintroduce yourself and sign-off References remote-work a guide to ensuring quality video conferencing ","date":"April 1, 2024","permalink":"/posts/technical_content_videos/","summary":"How to make techincal content videos.","title":"Technical Content Video Guidelines","type":"posts"},{"content":" Docs Structure Getting Started I have 5 min. Does this project work? All visitors find this information helpful.\nHigh-level intro to features Purpose and benefits Who should use the solution Developer/Technical level intro Introduce key features Introduce Bindings Describe what the solution does using generic terms Persona/Solution specific use cases I have a specific need and heard this project can help. Main visitor filter.\nSpecialized situations. Enterprise Product integrations Full guides Live Demos Explicit/Vanilla Configuration I have a working solution that requires more configuration or a use case that isn\u0026rsquo;t covered.\nDeep dive into the solution config Comprehensive list of use cases Clear descriptions and concise examples Quickstart Show me what your solution does.\nShort getting started Important places to continue learning Highlight the solution config making it happen Use it to point users to the next step At the end Download everything and run it locally Security I have a working PoC; how do I add security?\nTLS and encryption Securing APIs Authentication and authorization When and where to add these components Generic Examples Deploy and Operate How do I get my solution into production?\nLocal deployment and Development environments Production deploy examples. Cloud, on-prem, Marketplace Automation and orchestration. Helm, Ansible, Terraform Describe Telemetry Logging \u0026amp; Monitoring with best practice examples Health checks and alerting Scaling and performance tuning Cloud Provider walkthroughs Reference \u0026amp; Glossary I am implementing my solution and need to understand what a specific attribute is\nComprehensive overview Mention everything visible to the user Every attribute described with Reader friendly snippets Automatic Reference linking Programmatically maintain/generate Troubleshooting How can I solve a problem that isn\u0026rsquo;t described in the docs?\nCommon issues and solutions Uncommon options and settings to aid in Debugging GH issues and discussions callout Community and paid support Guides and Tutorials General structure\nInstall Components install easily Connect Portable client connection Success Measured success criteria Situational:\nTestable Make it your own Use your own client Bring your own components Quickstarts Cover as much core functionality as possible. Create everything or the user and let them see a result.\nLow barrier, fast success, 5min proof of function Used to confirm the solution does what it says hosted Cookbooks/Guides Common feature combinations.\nTeaching focused Organized by function/usage Exhaustive explanation Used to confirm the solution does work for more situations Variable components Feature Examples Exhausted list of simple and isolated examples.\nFeature focused Organized by category Short copy \u0026amp; paste oriented and testable Used to add or troubleshoot features Self contained Static components Use Case Demos Full end to end demos brining in other services or applications\nShow and Tell focused End to End, Real world scenarios Quick setup and walkthrough Used to showcase multiple functions Common Config Templates Assistive config generation.\nFill in the blank configs I want to bootstrap a solution I will bring my own components Docs Inspiration Logical Reference pkg.go.dev/net/http Practical Organization questdb.io/docs/ vuejs.org/guide/introduction.html docs.ansible.com/ developer.hashicorp.com/terraform/docs docs.adyen.com/development-resources/ infisical.com/docs/documentation/getting-started/introduction docs.redpanda.com/redpanda-connect/about/ Examples vuejs.org/examples/#hello-world getstream.io/chat/docs/sdk/android/compose/overview/ ","date":"November 1, 2023","permalink":"/posts/organizing_documentation/","summary":"How to Organize docs","title":"How to Organize docs","type":"posts"},{"content":"The single purpose of any content is to be consumed. As new topic discovery happens through observation of feeds, topic research happens through search first and category navigation second. Search optimization benefits the consumer, accurate categorization benefits the creator. All levels of content from documentation to blogs should be discovered through internal search tools and external search engine optimization.\nSearch optimization cannot lead to metadata content. A piece of content must first be created for the sole purpose of benefiting the consumer. Improved algorithmic results at the cost of quality or relevance will sour an audience\u0026rsquo;s attention.\nStakeholders Consumer This individual follows the documentation, reads articles, and/or watches videos. This needs to be identified for each piece of content in order to cater the massage to the audience.\nCreator An individual who spent time in order to share an idea with others. Will be captured to assist in updating and cataloging efforts.\nManager Those responsible for the organization of content. A dedicated team that focuses on organization and ease of access to content.\nAdmin Those responsible for distributing and hosting content. A more technical focus to work out how to implement content management.\nProperties A separate and siloed organization benefits content management instead of consumers. An effort needs to be made to collect all resources under the same domain. Additionally creating a more uniform delivery across different products and BUs.\nDocumentation - Truth This is the primary decision point. Any mistake here is projected onto the product. All other sources of content point a developer to this baseline in order to execute an idea.\nSpec Granular list and description for each and every component.\nFeature A concise breakdown of purpose and functionality.\nExecution Short single action breakdown.\nResources - Learn A purpose-built collection of articles meant to educate, remaining relevant well after the published date. Cataloged such that individual topics can be easily found. Categorization and search optimization are equally important to the material.\nHands-On How-tos, Guides, Samples Long-form description backed by code and/or documentation. Descriptions are hosted but often point to external platforms for execution. The organization of Git repositories will benefit consumer discovery.\nEvaluate White Papers, Thought Leadership, Stories, Reports{#white-papers,-thought-leadership,-stories,-reports} Content that will be relevant for longer than a blog and potentially needs to be updated. Discussing\nRoadmap Keeping sources for sharing product vision will allow managers to share with consumers. This can be presented as general goals and updated in place as items change.\nBlog - Discover A unified voice streaming ideas organized by topic and audience. A river delta celebrating the potential solution to a real or perceived problem. Reporting new information good and bad.\nCategorize All content should be discoverable in the same place. Filtering by subject and tags will allow consumers to cater to their own needs or interests.\nProducts Each product can have a dedicated stream but there will still be company-wide content that won\u0026rsquo;t need to fit under one specific product.\nAnnouncements Company news, changelogs, etc will benefit the consumer by being all in one place and in the same format.\nEcosystem - Connect Community Groups of consumers share interests and should be united wherever possible. While a managing team may want to have separate need groups a consumer would rather find everything in one place.\nEvents This requires a larger time commitment for the consumer and should be focused on a single purpose.\nForums A living history of public discussion that should be available to consume by anyone but needs to be moderated such that the discussion has a clear topic and resolution.\nPipeline Idea, Create, Publish A single process with few measured deviations. Capture, cultivate, and curate in order to improve quality and benefit the consumer.\nSubmission Having one place to start this process would be ideal, but the fewer entrances to a pipeline the better. With one place writers can submit a piece the direction and discovery can begin for any type of content. A single structure can be communicated and enforced. Any specific topic prompting can be listed.\nValidation Let tools work for us. Adding automatic text validation will increase quality without having to suggest writers use specific tools.\nFeedback The immediate benefit for creators. Asking someone to invest the time can be reworded by the automated and collaborative observation from the team. Even when content doesn\u0026rsquo;t fit into any need for the company the discussion and improvements will benefit the creator.\nChannels and statuses should cater to what teams can/should provide feedback or become aware of the new content.\nVisibility When content is validated the discussion needs to include more stakeholders. Increasing the internal visibility will improve the overall impact. Additionally, once content is published there needs to be multiple people watching for any further feedback from the consumers\nCatalog Putting content in the right place matters both to reach the correct audience and to ensure proper ongoing maintenance. A detailed and clear appendix of all relevant content will keep mistakes low and easier to fix. Stale can be just as harmful as wrong, knowing when and what to update or archive will be invaluable.\nPromotion Marketing teams want to find the right content as much as the consumers do. Deciding what type and area will give them a backlog to work with as well as help identify any gaps.\nMaintenance Publishing is not the end of the line. Dedicated efforts to ensure existing content remain relevant or is moved to an appropriate place enforce the importance of new content and reduce confusion.\nReview Quarterly meta-analysis of all content will encourage reuse and remove outdated content. Additionally making feedback channels for each piece of content can alert to problems early.\nPatch Not all content needs to be removed or replaced. The ideal case is to fix errors or update content in place. An article that has been edited can be shared as new or have a more recent published date.\nArchive Some content does more harm than good. Proper collection and storage in a single archive will give clarity to the consumer.\n","date":"November 1, 2022","permalink":"/posts/search_first_content/","summary":"All content that is meant for consumption should be easily searchable.","title":"Search First Content","type":"posts"},{"content":"The developer relations team should have its roots in engineering. Empathy for the audience they need to serve is essential. Advocacy is first for the developer community. This community is any developer involved in the product, internal or external, past and present. In a Software tool company, this means overlap in different departments. DevRel is the least biased path to a developer\u0026rsquo;s first success with any new tool or product.\nTo the developer community, represent the company. To the company, represent the community.\nFocus areas Documentation\nWindow to the product Trust is earned or lost Simple \u0026amp; Thorough Samples \u0026amp; SDKs\nModern Easy to Use Always Open Source Guides\nExpand on Documentation Are use case driven First Success\nHelping others Reduce friction Time to Value focus Foster Community\nChat groups Answer questions Meetups Write, Record, and Share Areas DevRel can Help Dev Experience Accurate/Searchable Resources Language libraries and SDKs Getting started guides Self-service/Open-Source Community chat \u0026amp; question forum Sample applications Developer blog posts Audit Feature/API references Service status/uptime info Interactive documentation Community Local events, Meetup/User groups Audio/Video content Social Media Ambassador program Swag store ","date":"September 1, 2022","permalink":"/posts/devrel_quality_over_vanity/","summary":"DevRel should be measured by the Quality of the experience.","title":"Effective DevRel means Quality over Vanity","type":"posts"},{"content":"The benefit of a standalone identity platform starts on the first day of planning your project and continues through the life of the application.\nOriginally published on The New Stack by Andrew Danelz A problem starts to reveal itself. The team has finished a minimum viable product\u0026rsquo;s (MVP) business features, and user personally identifiable information (PII) has merged into the same domain as other core objects. Social logins bring their own attributes that need to be captured but don\u0026rsquo;t line up. User privacy becomes a concern, and your app hasn\u0026rsquo;t requested consent to store all the user information.\nYour dev team starts to implement the advanced identity and access management (IAM) features but needs extra time to do proper research since these aren\u0026rsquo;t out-of-the-box capabilities in the development framework. The team has to pay back all of the tech debt taken to get the MVP out on time.\nThe user is at the center of any new project. The lowest-effort user implementation comes with limited capabilities. Building early features on top of a simple structure will lead to design choices that will require updates to those core structures.\nWhen stronger security, observability or improved user experience is a requirement later in the roadmap, there will be challenges adapting the existing identity systems. Disconnected user silos and workflows will all need to be updated.\nStarting with a robust identity and access management (IAM) solution will give new projects a head start on the competition. Users will have access to more features earlier. Additionally, no growing pains also mean no tech debt.\nAny new project has challenges right from the start. Finishing the MVP is a high priority. Planning meetings to outline necessary features and requirements can suffer from scope creep. Every shortcut taken to deliver on time borrows against the future. Tech debt is a known cost, and many startups take on a significant amount.\nAs any app needs users, it eventually will come down to planning the features and structures needed. Everyone is a user themselves, so it\u0026rsquo;s easy to come up with a variety of useful features. Single sign-on , social logins and multifactor authentication are all conventional IAM features included in the project scope and planned out for customers. Features and domain knowledge are designed around what the team thinks a user will need. A user\u0026rsquo;s footprint within your app gets built out in forms and user profile pages. Business data and user data are stored together.\nThe plan is set and now it is time to implement it. Your team is confident because the new framework has user authentication built-in. The samples are followed to the letter and the first working prototype is finished. Nothing fancy, but just the bare minimum. A user can log in with an email and password just fine. It has all those nice IAM features and requirements for onboarding and stacking up against the core business domain needs. The team can move to the backlog to be addressed when they are needed.\nThere is a better way!\nIntegrating a standalone identity platform requires slightly more effort upfront, but will save on tech debt later. The benefit to your project starts on the first day of planning and through the life of the application. All the user authentication stories are handled by the identity platform and only need to be configured. This will ensure you start the app with a clear understanding of a user\u0026rsquo;s privacy. Your app will need user consent to access PII, making certain that domain data doesn\u0026rsquo;t get mixed in. The application is built around a user identified by a unique ID that meets privacy guidelines.\nYour MVP will not be delayed, and your user will have the standard IAM features that your team didn\u0026rsquo;t have to write. When the time comes to improve the user experience with single sign-on, social logins and multifactor authentication, there is no tech debt to pay back. Your team doesn\u0026rsquo;t need to become experts, but rather follow clear documentation to add these features.\nWith a small investment early into a proven identity platform, your team will have a clear picture of how to interact with users, and your users will have more features on launch. Improving user experience happens easily as the need for increased security grows. Your users aren\u0026rsquo;t siloed, and additional applications can be integrated with ease with your existing user base.\n","date":"March 30, 2022","permalink":"/posts/iam_tech_debt/","summary":"The benefit of a standalone identity platform starts on the first day of planning your project and continues through the life of the application.","title":"With Identity Management, Start Early for Less Tech Debt","type":"posts"},{"content":"The link below goes to Dan McKinley\u0026rsquo;s talk Egoless Engineering. The content is great but the format makes it hard to share, so here is a written summary of the argument and the takeaways.\nMcKinley\u0026rsquo;s core claim is that most organizational dysfunction in engineering comes from how companies divide responsibility. Roles get split into ever-narrower specialties, ownership turns into territory, and the boundaries between teams generate work that nobody wanted. Cooperation is the antidote, and leaders are the ones who make it possible or impossible.\nA few examples he uses to make the point:\nA startup needed database specialists, Python developers, and PHP developers to coordinate on every feature. They shipped nothing for two years. A mature company kept subdividing roles until it invented a \u0026ldquo;release manager\u0026rdquo; position. The role became a bottleneck. A company assigned all security remediation to a dedicated security team to reduce context switching for engineers. Developers stopped thinking about security at all, which created more work, not less. He also gives a positive example. At one company, a designer broke the build. Rather than putting up walls between designers and the deploy pipeline, leadership gave the designer deploy keys. After two years of similar moves, almost everyone contributed to almost everything. Domain experts mentored. They did not gatekeep.\nKey takeaways How you divide responsibility shapes everything. The choice has immediate effects, and worse second-order effects. Poorly factored boundaries create work. Handoffs, queues, and approvals all show up as labor the org pays for. Ego and parochialism are the failure modes. As McKinley puts it, the biggest problem with brilliant jerks is not that they are jerks. They are not even brilliant. Misery is a bad proxy for results. Mandatory process, ceremony, and busy-ness are not the same as output. Domain experts, not domain owners. Specialists multiply their impact by teaching, not by hoarding tickets. Expand the in-group on purpose. Bootcamps, hack weeks, and forced cross-functional work build the trust that makes cooperation cheap. Articulate anti-elitism out loud. Leaders have to say that no role is more technical or more valuable than another, then act like it. Protect organizational slack. Systems run at 100% utilization cannot learn, experiment, or cooperate. Slack is a feature. Leaders set the tone. Most leaders never cut feel-bad programs like mandatory review, but they cut feel-good programs like hack weeks first. Notice when you are doing that. Cooperation needs permission. People wait to be authorized to be curious or to help across team boundaries. Give the permission explicitly. Links egoless.engineering — Dan McKinley\u0026rsquo;s original talk. ","date":"February 1, 2022","permalink":"/posts/egoless_engineering/","summary":"Dan McKinley\u0026rsquo;s talk Egoless Engineering","title":"A summary of Egoless Engineering","type":"posts"},{"content":"Some good ideas pull from https://plane.so/blog/how-we-got-to-20k-github-stars to discuss and add to.\n1: Community What worked #1\nWe crowdsourced ideas for comms, not just the product. We asked them to outline their business goals, describe their troubles with other software, name our features, even opine on the logo. We consistently kept them updated about our progress. What didn\u0026rsquo;t #1\nAutomating updates and boilerplate responses Delays with Support requests, especially from new users 2: Release notes What worked #2\nActive comms, probing for the whys behind every request, and workarounds Personal messages when a feature got shipped Asking for referrals What didn\u0026rsquo;t #2\nLengthy messages with unnecessary context Shying away from sharing work in progress Not creating a dedicated forum for Support early on In hindsight, this was the game prep we needed before we went wide with our repo and promotions. It let us see challenges ahead of time, plan for them, and execute with intent.\n3: Repurposing content What worked #3\nIt didn\u0026rsquo;t push Plane into anyone\u0026rsquo;s face. Short-and-sweet instructions for self-hosting came first. There were links to destinations new users expect to see. Using simple language, keeping in mind of the global audience What didn\u0026rsquo;t #3\nThe full readme is a little too long. We heard some folks saying it was too promotional—something we are still considering changing. 4: Deliberate distribution What worked #4\nCycling posts through five destinations Bringing older updates back up to the top of our users\u0026rsquo; feeds Figuring out destinations and networks that were working well Doubling down on them and discarding others in the short-term What didn\u0026rsquo;t #4\nNo variation in language or tone for different channels What worked #5\nFocusing on one channel before adding any other Experimentation with formats, tones, and styles until we found our Reddit sweet spot Direct access to a release on GitHub from its Reddit post 5: Captive audiences What didn\u0026rsquo;t #5\nTalking about every bug or minor improvement Excessive promotion or cross-posting 6: Hacking Hacker News What worked #6\nAlways tagging contributors and other open-source projects Linking metrics for releases when announcing them Showing early previews of features in action, usually via a GIF What didn\u0026rsquo;t #6\nAutomating tweets without careful, careful planning Linking to our sites and resources a little too much without enough show-and-tell sometimes Tactic #7: Crisis comms\nWhat worked #7\nStrategic publisher and audience research Enough breadcrumbs to the product, Docs, or the plane.so site Lots of show-don\u0026rsquo;t-tell What didn\u0026rsquo;t #7\nNot planning for reviews and approvals Going wide, not deep, with distribution Almost-zero calls to action ","date":"November 1, 2021","permalink":"/posts/getting_github_stars/","summary":"Ideas on how to get more GitHub stars","title":"Getting GitHub Stars","type":"posts"},{"content":"Source: github.com/tsiege/Tech-Interview-Cheat-Sheet This list is meant to be both a quick guide and reference for further research into these topics. It\u0026rsquo;s basically a summary of that comp sci course you never took or forgot about, so there\u0026rsquo;s no way it can cover everything in depth.\nTable of Content Asymptotic Notation Complexities Big-O notation Big-Ω (Big-Omega) notation Big-θ (Big-Theta) notation What you need to know Data Structures Array What you need to know Time Complexity Linked List What you need to know Time Complexity Hash Table or Hash Map What you need to know Time Complexity Binary Tree What you need to know Time Complexity Algorithm Basics Recursive Algorithms What you need to know Iterative Algorithms What you need to know Recursion Vs. Iteration Pseudo Code of Moving Through an Array Greedy Algorithms What you need to know Pseudo Code of a Greedy Algorithm to Find Largest Difference of any Two Numbers in an Array. Search Algorithms Breadth First Search What you need to know Time Complexity Depth First Search What you need to know Time Complexity Breadth First Search Vs. Depth First Search Nuances Sorting Algorithms Selection Sort What you need to know Time Complexity Space Complexity Visualization Insertion Sort What you need to know Time Complexity Space Complexity Visualization Merge Sort What you need to know Time Complexity Space Complexity Visualization Quicksort What you need to know Time Complexity Space Complexity Visualization Merge Sort Vs. Quicksort Additional Resources Asymptotic Notation Asymptotic Notation is the hardware independent notation used to tell the time and space complexity of an algorithm. Meaning it\u0026rsquo;s a standardized way of measuring how much memory an algorithm uses or how long it runs for given an input.\nComplexities The following are the Asymptotic rates of growth from best to worst:\nconstant growth - O(1) Runtime is constant and does not grow with n logarithmic growth – O(log n) Runtime grows logarithmically in proportion to n linear growth – O(n) Runtime grows directly in proportion to n superlinear growth – O(n log n) Runtime grows in proportion and logarithmically to n polynomial growth – O(n^c) Runtime grows quicker than previous all based on n exponential growth – O(c^n) Runtime grows even faster than polynomial growth based on n factorial growth – O(n!) Runtime grows the fastest and becomes quickly unusable for even small values of n (source: Soumyadeep Debnath, Analysis of Algorithms | Big-O analysis) Visualized below; the x-axis representing input size and the y-axis representing complexity:\n(source: Wikipedia, Computational Complexity of Mathematical Operations) Big-O notation Big-O refers to the upper bound of time or space complexity of an algorithm, meaning it worst case runtime scenario. An easy way to think of it is that runtime could be better than Big-O but it will never be worse.\nBig-Ω (Big-Omega) notation Big-Omega refers to the lower bound of time or space complexity of an algorithm, meaning it is the best runtime scenario. Or runtime could worse than Big-Omega, but it will never be better.\nBig-θ (Big-Theta) notation Big-Theta refers to the tight bound of time or space complexity of an algorithm. Another way to think of it is the intersection of Big-O and Big-Omega, or more simply runtime is guaranteed to be a given complexity, such as n log n.\nWhat you need to know Big-O and Big-Theta are the most common and helpful notations Big-O does not mean Worst Case Scenario, Big-Theta does not mean average case, and Big-Omega does not mean Best Case Scenario. They only connote the algorithm\u0026rsquo;s performance for a particular scenario, and all three can be used for any scenario. Worst Case means given an unideal input, Average Case means given a typical input, Best case means a ideal input. Ex. Worst case means given an input the algorithm performs particularly bad, or best case an already sorted array for a sorting algorithm. Best Case and Big Omega are generally not helpful since Best Cases are rare in the real world and lower bound might be very different than an upper bound. Big-O isn\u0026rsquo;t everything. On paper merge sort is faster than quick sort, but in practice quick sort is superior. Data Structures Array Stores data elements based on an sequential, most commonly 0 based, index. Based on tuples from set theory. They are one of the oldest, most commonly used data structures. What you need to know Optimal for indexing; bad at searching, inserting, and deleting (except at the end). Linear arrays, or one dimensional arrays, are the most basic. Are static in size, meaning that they are declared with a fixed size. Dynamic arrays are like one dimensional arrays, but have reserved space for additional elements. If a dynamic array is full, it copies its contents to a larger array. Multi dimensional arrays nested arrays that allow for multiple dimensions such as an array of arrays providing a 2 dimensional spacial representation via x, y coordinates. Time Complexity Indexing: Linear array: O(1), Dynamic array: O(1) Search: Linear array: O(n), Dynamic array: O(n) Optimized Search: Linear array: O(log n), Dynamic array: O(log n) Insertion: Linear array: n/a, Dynamic array: O(n) Linked List Stores data with nodes that point to other nodes. Nodes, at its most basic it has one datum and one reference (another node). A linked list chains nodes together by pointing one node\u0026rsquo;s reference towards another node. What you need to know Designed to optimize insertion and deletion, slow at indexing and searching. Doubly linked list has nodes that also reference the previous node. Circularly linked list is simple linked list whose tail, the last node, references the head, the first node. Stack, commonly implemented with linked lists but can be made from arrays too. Stacks are last in, first out (LIFO) data structures. Made with a linked list by having the head be the only place for insertion and removal. Queues, too can be implemented with a linked list or an array. Queues are a first in, first out (FIFO) data structure. Made with a linked list that only removes from head and adds to tail. Time Complexity Indexing: Linked Lists: O(n) Search: Linked Lists: O(n) Optimized Search: Linked Lists: O(n) Append: Linked Lists: O(1) Prepend: Linked Lists: O(1) Insertion: Linked Lists: O(n) Hash Table or Hash Map Stores data with key value pairs. Hash functions accept a key and return an output unique only to that specific key. This is known as hashing, which is the concept that an input and an output have a one-to-one correspondence to map information. Hash functions return a unique address in memory for that data. What you need to know Designed to optimize searching, insertion, and deletion. Hash collisions are when a hash function returns the same output for two distinct inputs. All hash functions have this problem. This is often accommodated for by having the hash tables be very large. Hashes are important for associative arrays and database indexing. Time Complexity Indexing: Hash Tables: O(1) Search: Hash Tables: O(1) Insertion: Hash Tables: O(1) Binary Tree Is a tree like data structure where every node has at most two children. There is one left and right child node. For a full binary-tree reference see Here What you need to know Designed to optimize searching and sorting. A degenerate tree is an unbalanced tree, which if entirely one-sided, is essentially a linked list. They are comparably simple to implement than other data structures. Used to make binary search trees. A binary tree that uses comparable keys to assign which direction a child is. Left child has a key smaller than its parent node. Right child has a key greater than its parent node. There can be no duplicate node. Because of the above it is more likely to be used as a data structure than a binary tree. Time Complexity Indexing: Binary Search Tree: O(log n) Search: Binary Search Tree: O(log n) Insertion: Binary Search Tree: O(log n) Algorithms Algorithm Basics Recursive Algorithms An algorithm that calls itself in its definition. Recursive case a conditional statement that is used to trigger the recursion. Base case a conditional statement that is used to break the recursion. What you need to know Stack level too deep and stack overflow. If you\u0026rsquo;ve seen either of these from a recursive algorithm, you messed up. It means that your base case was never triggered because it was faulty or the problem was so massive you ran out of alloted memory. Knowing whether or not you will reach a base case is integral to correctly using recursion. Often used in Depth First Search Iterative Algorithms An algorithm that is called repeatedly but for a finite number of times, each time being a single iteration. Often used to move incrementally through a data set. What you need to know Generally you will see iteration as loops, for, while, and until statements. Think of iteration as moving one at a time through a set. Often used to move through an array. Recursion Vs. Iteration The differences between recursion and iteration can be confusing to distinguish since both can be used to implement the other. But know that, Recursion is, usually, more expressive and easier to implement. Iteration uses less memory. Functional languages tend to use recursion. (i.e. Haskell) Imperative languages tend to use iteration. (i.e. Ruby) Check out this Stack Overflow post for more info. Pseudo Code of Moving Through an Array Recursion | Iteration ----------------------------------|---------------------------------- recursive method (array, n) | iterative method (array) if array[n] is not nil | for n from 0 to size of array print array[n] | print(array[n]) recursive method(array, n+1) | else | exit loop | Greedy Algorithms An algorithm that, while executing, selects only the information that meets a certain criteria. The general five components, taken from Wikipedia : A candidate set, from which a solution is created. A selection function, which chooses the best candidate to be added to the solution. A feasibility function, that is used to determine if a candidate can be used to contribute to a solution. An objective function, which assigns a value to a solution, or a partial solution. A solution function, which will indicate when we have discovered a complete solution. What you need to know Used to find the expedient, though non-optimal, solution for a given problem. Generally used on sets of data where only a small proportion of the information evaluated meets the desired result. Often a greedy algorithm can help reduce the Big O of an algorithm. Pseudo Code of a Greedy Algorithm to Find Largest Difference of any Two Numbers in an Array. greedy algorithm (array) var largest difference = 0 var new difference = find next difference (array[n], array[n+1]) largest difference = new difference if new difference is \u0026gt; largest difference repeat above two steps until all differences have been found return largest difference This algorithm never needed to compare all the differences to one another, saving it an entire iteration.\nSearch Algorithms Breadth First Search An algorithm that searches a tree (or graph) by searching levels of the tree first, starting at the root. It finds every node on the same level, most often moving left to right. While doing this it tracks the children nodes of the nodes on the current level. When finished examining a level it moves to the left most node on the next level. The bottom-right most node is evaluated last (the node that is deepest and is farthest right of it\u0026rsquo;s level). What you need to know Optimal for searching a tree that is wider than it is deep. Uses a queue to store information about the tree while it traverses a tree. Because it uses a queue it is more memory intensive than depth first search. The queue uses more memory because it needs to stores pointers Time Complexity Search: Breadth First Search: O(V + E) E is number of edges V is number of vertices Depth First Search An algorithm that searches a tree (or graph) by searching depth of the tree first, starting at the root. It traverses left down a tree until it cannot go further. Once it reaches the end of a branch it traverses back up trying the right child of nodes on that branch, and if possible left from the right children. When finished examining a branch it moves to the node right of the root then tries to go left on all it\u0026rsquo;s children until it reaches the bottom. The right most node is evaluated last (the node that is right of all it\u0026rsquo;s ancestors). What you need to know Optimal for searching a tree that is deeper than it is wide. Uses a stack to push nodes onto. Because a stack is LIFO it does not need to keep track of the nodes pointers and is therefore less memory intensive than breadth first search. Once it cannot go further left it begins evaluating the stack. Time Complexity Search: Depth First Search: O(|E| + |V|) E is number of edges V is number of vertices Breadth First Search Vs. Depth First Search The simple answer to this question is that it depends on the size and shape of the tree. For wide, shallow trees use Breadth First Search For deep, narrow trees use Depth First Search Nuances Because BFS uses queues to store information about the nodes and its children, it could use more memory than is available on your computer. (But you probably won\u0026rsquo;t have to worry about this.) If using a DFS on a tree that is very deep you might go unnecessarily deep in the search. See xkcd for more information. Breadth First Search tends to be a looping algorithm. Depth First Search tends to be a recursive algorithm. Sorting Algorithms Selection Sort A comparison based sorting algorithm. Starts with the cursor on the left, iterating left to right Compares the left side to the right, looking for the smallest known item If the left is smaller than the item to the right it continues iterating If the left is bigger than the item to the right, the item on the right becomes the known smallest number Once it has checked all items, it moves the known smallest to the cursor and advances the cursor to the right and starts over As the algorithm processes the data set, it builds a fully sorted left side of the data until the entire data set is sorted Changes the array in place. What you need to know Inefficient for large data sets. Very simple to implement. Time Complexity Best Case Sort: O(n^2) Average Case Sort: O(n^2) Worst Case Sort: O(n^2) Space Complexity Worst Case: O(1) Visualization (source: Wikipedia, Selection Sort) Insertion Sort A comparison based sorting algorithm. Iterates left to right comparing the current cursor to the previous item. If the cursor is smaller than the item on the left it swaps positions and the cursor compares itself again to the left hand side until it is put in its sorted position. As the algorithm processes the data set, the left side becomes increasingly sorted until it is fully sorted. Changes the array in place. What you need to know Inefficient for large data sets, but can be faster for than other algorithms for small ones. Although it has an O(n^2) time complexity, in practice it is slightly less since its comparison scheme only requires checking place if it is smaller than its neighbor. Time Complexity Best Case: O(n) Average Case: O(n^2) Worst Case: O(n^2) Space Complexity Worst Case: O(n) Visualization (source: Wikipedia, Insertion Sort) Merge Sort A divide and conquer algorithm. Recursively divides entire array by half into subsets until the subset is one, the base case. Once the base case is reached results are returned and sorted ascending left to right. Recursive calls are returned and the sorts double in size until the entire array is sorted. What you need to know This is one of the fundamental sorting algorithms. Know that it divides all the data into as small possible sets then compares them. Time Complexity Worst Case: O(n log n) Average Case: O(n log n) Best Case: O(n) Space Complexity Worst Case: O(1) Visualization (source: Wikipedia, Merge Sort) Quicksort A divide and conquer algorithm Partitions entire data set in half by selecting a random pivot element and putting all smaller elements to the left of the element and larger ones to the right. It repeats this process on the left side until it is comparing only two elements at which point the left side is sorted. When the left side is finished sorting it performs the same operation on the right side. Computer architecture favors the quicksort process. Changes the array in place. What you need to know While it has the same Big O as (or worse in some cases) many other sorting algorithms it is often faster in practice than many other sorting algorithms, such as merge sort. Time Complexity Worst Case: O(n^2) Average Case: O(n log n) Best Case: O(n log n) Space Complexity Worst Case: O(log n) Visualization (source: Wikipedia, Quicksort) Merge Sort Vs. Quicksort Quicksort is likely faster in practice, but merge sort is faster on paper. Merge Sort divides the set into the smallest possible groups immediately then reconstructs the incrementally as it sorts the groupings. Quicksort continually partitions the data set by a pivot, until the set is recursively sorted. Additional Resources Khan Academy\u0026rsquo;s Algorithm Course Graph Data Structure \u0026amp; Algorithms Data Structure Interview Questions Data Structure MCQ With Answers 10 Best Data Structures and Algorithms Books ","date":"December 1, 2020","permalink":"/posts/interview_study/","summary":"This list is meant to be both a quick guide and reference for further research into these topics. It\u0026rsquo;s basically a summary of that comp sci course you never took or forgot about, so there\u0026rsquo;s no way it can cover everything in depth.","title":"Tech Interview Cheat Sheet","type":"posts"},{"content":" The Mikado Method The Pragmatic Programmer Modern Software Engineering Code Complete: A Practical Handbook of Software Construction Software Engineering at Google Clean Code Head First Design Patterns Refactoring Grokking Algorithms Introduction to Algorithms Designing Data-Intensive Applications Learning SQL Growing OO Software by Tests Unit Testing Principles, Practices, and Patterns The Art of Unit Testing TDD by Example Fundamentals Of Software Architecture Clean Architecture Software Architecture The Hard Parts Domain Driven Design Quickly A Philosophy of Software Design Understanding Distributed Systems Designing Distributed Systems DevOps Handbook Continuous Delivery Accelerate The Hundred-Page Machine Learning Book Designing Machine Learning Systems The Five Dysfunctions of a Team Drive The Making of a Manager How to Win Friends \u0026amp; Influence People Deep Work The 7 Habits of Highly Effective People Atomic Habits A Programmer\u0026rsquo;s Reading List: 100 Articles I Enjoyed (1-50) Software Engineering and Career 📚 Books 📚 Clean Code 📚 The clean coder 📚 A Philosophy of Software Design 📚 The Algorithm Design Manual 📚 Dive into design patterns (refactoring.guru ) 📚 Design Patterns: Elements of Reusable Object-Oriented Software 📚 Cracking the Coding Interview 📚 Grokking Algorithms 📚 Introduction to Algorithms 📚 Effective Java 📚 Software Engineering at Google 📚 The Complete Software Development Career Guide 📚 Modern Software Engineering 📚 The Missing README 📚 The Pragmatic Programmer 📚 The Software Engineer\u0026rsquo;s Guidebook 📚 The Effective Engineer System Design and Architecture 📚 Books 📚 Cracking the coding interview 📚 Web Scalability for Startup Engineers 📚 Building Microservices by Sam Newman 📚 Monolith to Microservices by Sam Newman 📚 Designing Data Intensive Applications 📚 Understanding Distributed Systems 📚 Fundamentals of Software Architecture: An Engineering Approach 📚 Software Architecture “The Hard Parts” 📚 System Design Interview book 1 and book 2 Code Complete by Steve McConnell\nDevOps Handbook\n","date":"January 1, 2012","permalink":"/posts/reading_list/","summary":"A page maintaining my reading list","title":"My Tech Reading List","type":"posts"},{"content":" My name is AJ Danelz Software Engineer and Advocate I am a full-stack cloud application engineer and developer advocate passionate about distributed cloud applications. I am a golang enthusiast trapped in a Java dev\u0026rsquo;s body. I have been a cloud-native aficionado since the moment I learned there was a world outside of VMs and server racks. My passions for event streaming and microservices fit nicely in the cloud. I love working for a global company, working with both technical and non-technical team members. I have served in a leadership role for the majority of my development career. In any role, I will contribute to the success of a software project by advocating internally and externally for all developers.\nPortfolio | linkt.ree Skills DISCLAIMER I have a personal policy against any type of live coding or online coding tests during interviews. Otherwise, I am happy to work offline on coding assignments with reasonable goals and deadlines and to have in-depth technical discussions about software architecture and design as well as relevant technologies.\n","permalink":"/profile/","summary":"My name is AJ Danelz Software Engineer and Advocate I am a full-stack cloud application engineer and developer advocate passionate about distributed cloud applications. I am a golang enthusiast trapped in a Java dev\u0026rsquo;s body. I have been a cloud-native aficionado since the moment I learned there was a world outside of VMs and server racks. My passions for event streaming and microservices fit nicely in the cloud. I love working for a global company, working with both technical and non-technical team members.","title":"","type":"page"},{"content":" Roasting section instructions This section logs coffee roasts on a Behmor 2000 AB Plus. The architecture separates universal roast science from machine-specific guidance so future machines can be added without template changes.\nFile layout content/roasts/00N-\u0026lt;origin\u0026gt;.md — one file per roast. Structured frontmatter holds plan + raw inputs + observations; the body is rendered from those fields, not from markdown prose. content/roasts/_TEMPLATE.md — scaffold for new roasts. Excluded from Tina collection via match.exclude: \u0026quot;{_*,methods,behmor-*}\u0026quot;. content/roasts/methods.md — universal-only reference doc. Formulas, target-level ranges, raw input definitions. No machine names here. content/roasts/behmor-2000ab.md — Behmor 2000 AB Plus machine cheat sheet. Tables render from YAML via shortcodes — do not hand-edit the tables; edit the data file. data/roast_guidance.yaml — universal: levels (loss%/DTR%/rest_days/drop/flavor per roast level), formulas. No machine references. data/machines/\u0026lt;model_key\u0026gt;.yaml — everything machine-specific: profiles, weight_settings, preheat, safety, temperature_buttons, drum_speed. Has a reference_url field pointing to its cheat-sheet page. Machine selection is data-driven layouts/roasts/single.html derives the machine YAML key from the roaster frontmatter field:\n\u0026#34;Behmor 2000 AB Plus\u0026#34; → lower → replaceRE `[^a-z0-9]+` \u0026#34;_\u0026#34; → \u0026#34;behmor_2000_ab_plus\u0026#34; That key looks up .Site.Data.machines.behmor_2000_ab_plus. If no match, the machine panel silently omits machine-specific data (this is intentional — the three earliest roasts still carry the wrong-model roaster: Behmor 1600 Plus value).\nAdding a new machine Create data/machines/\u0026lt;model_key\u0026gt;.yaml matching the shape of behmor_2000_ab_plus.yaml (top-level: name, reference_url, profiles, weight_settings, preheat, safety, temperature_buttons, drum_speed). Optionally create content/roasts/\u0026lt;model-slug\u0026gt;.md as a human-readable cheat sheet. Use the existing roast-behmor-weights / roast-behmor-profiles shortcodes with machine=\u0026quot;\u0026lt;model_key\u0026gt;\u0026quot; arg. Add the new cheat-sheet link to the references box in layouts/roasts/list.html (hardcoded there, not in markdown). Future roasts use the new model name in roaster. No template or schema changes required. Calibrating ranges from logged data After ~10–15 logged roasts, the universal targets in data/roast_guidance.yaml may need tightening to match real cup outcomes. Edit the YAML — do not embed ranges in templates. All pages (single, list cards, methods, machine refs) pick up changes on rebuild.\nReady-to-drink window Each roast page and list card computes a rest window from the roast date + the level\u0026rsquo;s rest_days range, then shows a status pill (Resting / Ready / Past peak) by comparing against now. The math lives in layouts/roasts/single.html and layouts/roasts/list.html and uses time.ParseDuration with hour-based durations (mul 24 days → %dh). To tune rest windows, edit the rest_days arrays in data/roast_guidance.yaml — both pages update on rebuild.\nRest-time assumptions: the YAML ranges are tuned for drum roaster + filter brewing, which is what this site logs. Scott Rao: drum roasts don\u0026rsquo;t need more than 1–2 days unless underdeveloped; air roasters need 1–4 weeks. Filter vents CO₂ freely (vendor \u0026ldquo;12–24hr\u0026rdquo; minimums are real); espresso needs ~5–7 more days because gas can\u0026rsquo;t escape under 9 bar. Don\u0026rsquo;t widen these ranges back toward espresso/air-roaster numbers without changing the brew-method assumption documented in methods.md and the YAML header.\nRoasted beans photo roasted_photo is a Tina image field. Uploads go to static/imgs/ (configured in tina/config.ts media block). Phone uploads work because the project is on Tina Cloud. The single page renders the photo as a \u0026lt;figure\u0026gt; right after the headline stats strip, before the plan panel. Capped at 720px display width with max-width: min(100%, 720px) — responsive on phones, doesn\u0026rsquo;t dominate on desktop. No multi-resolution srcset yet — to add it later, move the Tina media root from static/ to assets/ and use Hugo\u0026rsquo;s image pipeline (resources.GetMatch + .Resize).\nBean source product_url (paste the vendor\u0026rsquo;s product page link) and optional vendor (human-readable name). The single page renders a \u0026ldquo;Source\u0026rdquo; row in the plan panel — uses vendor if set, otherwise extracts the URL\u0026rsquo;s host via replaceRE. Links open in a new tab with rel=\u0026quot;noopener\u0026quot;.\nBuilding a roast playbook — REQUIRED PROCESS A roast playbook directly drives a physical action on a $400 machine using $20+ of green coffee. Wrong guidance burns beans and risks fire (the Behmor manual is explicit about ignition past 10 seconds into 2C). Treat every playbook like instructions someone will follow with their hands on hot equipment.\nBefore writing any playbook step Read the machine reference page first — content/roasts/behmor-2000ab.md is the source of truth for everything machine-specific. The official manual is linked there.\nLook up the origin\u0026rsquo;s recommended Auto Mode profile from the machine YAML (auto_mode_profiles). The mapping is non-obvious:\nCentrals, Peruvian, Colombian → P1 or P2 (Hard Bean, highest heat) Brazilian, African, SE Asian, Malabar, JBM, Yauco → P3 (Soft Bean) Kona, low-grown island → P4 or P5 (lowest heat) Picking P5 for a Brazilian or Central American bean is wrong — P5 in Auto mode is the lowest heat profile, intended for Hawaiian/island coffees. The \u0026ldquo;P5 = aggressive\u0026rdquo; assumption is only correct in Manual mode (which fires AFTER Start).\nDefault to the C-button + Auto P1 hold technique for drinking-coffee batches at ½ lb or 1 lb: start on Auto P1 (or per-origin Auto profile), stay on Auto P1 through FC, drop to P3 only as a brake if 2C threatens. See c_button_technique in data/machines/behmor_2000_ab_plus.yaml.\nCRITICAL — the C button RESETS the clock, it does NOT add time. Pressing C sets the clock to the Rosetta Stone 1C→2C duration for the active weight (~1:30 for ¼ lb, ~2:09 for ½ lb, ~3:00+ for 1 lb). If the current clock has MORE time than that, C SHORTENS the roast. This bit us on roast 008: FC arrived at 7:30 with 4:30 remaining; pressing C at FC cut the clock to 2:09 and the bean only got 3:00 of post-FC development → 8.85% loss vs 16–18% target. Playbook rule: press C at FC ONLY if remaining clock \u0026lt; ~2:30. Otherwise let Auto P1\u0026rsquo;s original program run; use + for small extensions, or press C near end of clock if you need to extend.\nDo NOT default to the manual\u0026rsquo;s Pro Sample Roasting technique (P3+D at FC+15s) for drinking-coffee batches — it cuts power to 50% for the entire development phase, which stalls heat needed for moisture drive-off. Roasts 006 and 007 both used Pro Sample Roasting and came in under-developed (8–9.4% loss vs 13–16% target, DTR 38%+ vs 17–22% target). Pro Sample Roasting is for sample roasting (small evaluation batches where slow, low-development finish is intentional); see pro_sample_roasting.appropriate_use in the YAML.\nConfirm constraints from the manual:\nBatch ≤ 336 g if pushing past City+ (manual recommendation for darker control) 1:30 maximum preheat (longer locks out the roast) 1 HOUR between back-to-back roasts (NOT minutes — early playbooks got this wrong by an order of magnitude) Never past 10 sec into 2C 75% Err 7 safety shutoff fires at 75% of program time (need START or C to continue) Structure every playbook with A plan rationale block at the top — what bean type → which Auto profile, what target level, what charge size, expected FC behavior A numbered sequence that includes: preheat, drum insert, profile + weight + START, when to listen for FC, exact 75% shutoff time, C-button at FC (the default per c_button_technique in the machine YAML), drop trigger for the target level, cool cycle, weigh, and between-roast wait A \u0026ldquo;what to watch\u0026rdquo; block flagging the bean\u0026rsquo;s idiosyncrasies plus the universal fire-risk warnings (heavy smoke, past-2C, etc.) Cite the manual when adding non-obvious advice (e.g. \u0026ldquo;manual Part 5, page 12-13\u0026rdquo; for the pro technique). The citations are what let a future reader verify the playbook against an authoritative source. What NOT to do when building a playbook Don\u0026rsquo;t pick a profile from intuition — look it up in auto_mode_profiles. Don\u0026rsquo;t assume \u0026ldquo;more aggressive = higher P number\u0026rdquo; — that\u0026rsquo;s Manual mode logic, not Auto mode. Don\u0026rsquo;t write a back-to-back cool-down in minutes if the manual says 1 hour. Don\u0026rsquo;t reuse a prior roast\u0026rsquo;s playbook structure without checking it against the manual — early playbooks (001–004) were built before we had the manual extracted and contain known errors. Don\u0026rsquo;t copy generic drum-roaster advice (Genuine Origin, Scott Rao, etc.) into a Behmor playbook unless you\u0026rsquo;ve verified it applies to a probeless fixed-program machine. Don\u0026rsquo;t write unsourced sensory recipes (\u0026ldquo;press P3 for the final 30s before 2C\u0026rdquo;, \u0026ldquo;open door at 1:30 into cool\u0026rdquo;). If a step isn\u0026rsquo;t traceable to the manual or a cited forum/library, either drop it or mark it explicitly as common roasting knowledge / community technique. The rule is: a step under time pressure should not be dressed up as official guidance when it isn\u0026rsquo;t. Don\u0026rsquo;t test multiple variables in a single roast. When introducing a new technique (e.g. switching from Pro Sample Roasting to C-button + Auto P1 hold), validate it at a previously-attempted target before pushing to a new deeper or lighter level. Otherwise you can\u0026rsquo;t cleanly attribute an off result to technique vs target. If the user requests a multi-variable run anyway, flag it explicitly in bean_notes so the outcome is interpreted with that caveat. Every playbook runs in isolation Each NNN-*.md playbook must be written as if no other roast preceded it. Two specific consequences:\nAlways include the 1:30 preheat step. The Behmor\u0026rsquo;s cool cycle blows forced air through the chamber specifically to cool the drum and elements down. The next roast needs the standard preheat regardless of how recently the previous one finished. Do not write conditional preheat (\u0026ldquo;skip if back-to-back\u0026rdquo;) or \u0026ldquo;drum is hot\u0026rdquo; framing — both are wrong. No cross-references to other roasts in the playbook body. No \u0026ldquo;back-to-back\u0026rdquo; deviation framing, no \u0026ldquo;previous roast\u0026rdquo; notes, no chronology-based instructions. The 1-hour-between-roasts rule is a standalone instruction at the end of the sequence, not a framing for skipping steps. Bean notes may compare against prior tasting outcomes; the playbook itself cannot. The \u0026ldquo;no preheat, drum hot\u0026rdquo; pattern in roasts 003–004 is documented as one of the pre-manual errors — don\u0026rsquo;t propagate it.\nRepeatability — capture every button press The interventions field is the single most important repeatability tool. Starting settings (profile / weight / batch) are not enough to reproduce a roast — what makes the roast is the sequence of mid-roast button presses with their elapsed times. Capture them all:\nThe Manual-mode P switch (e.g. 14:00 — P3 (drop to 50%)) The D button if pressed (e.g. 14:00 — D (drum high)) Every C press (note: C is the documented extension lever for the C-button technique; resets the timer to the Rosetta Stone 1C→2C window) Every + press with the seconds added (note: + has a per-roast cumulative cap, ~+2:00 observed; if you reach for + as the primary extension lever you\u0026rsquo;ll run out) The Cool press at drop Format is one per line, MM:SS — what you pressed. Free text after the dash is fine. The page renders it as a monospace block in the Live notes section.\nWithout this log, \u0026ldquo;ran on Auto P1, 1 lb, dropped at 18:45\u0026rdquo; tells you nothing about when you switched modes, when you added time, or how the cup got where it did. With it, future-you can replicate a good roast or diagnose a bad one.\nCapturing machine details honestly When pulling info from the Behmor manual into the YAML data:\nQuote the manual directly in a manual_quote or similar field when the wording matters Flag contradictions rather than silently picking one. The manual gives two different drum-speed RPM ranges on different pages; the YAML now states this and treats the absolute values as approximate Don\u0026rsquo;t invent precision. If the manual says \u0026ldquo;high\u0026rdquo; and \u0026ldquo;standard\u0026rdquo; without RPM, capture those names. Don\u0026rsquo;t backfill numbers from forum posts or other machines This same standard applies to playbook construction — every non-obvious instruction should be traceable to a manual section. See \u0026ldquo;Building a roast playbook\u0026rdquo; above.\nTime fields time_to_fc and total_time are elapsed time in mm:ss. The Behmor displays countdown time — recording the countdown values produces a DTR \u0026gt; 40%, which is the immediate signal that elapsed vs countdown got confused. The first batch of roasts hit this; field descriptions in tina/config.ts clarify.\nWhat NOT to do Don\u0026rsquo;t embed roast-level ranges (loss %, DTR %) in the Hugo template — they belong in data/roast_guidance.yaml. Don\u0026rsquo;t hand-roll tables in methods.md or behmor-2000ab.md for data that lives in YAML — use the shortcodes (roast-formulas, roast-levels, roast-behmor-weights, roast-behmor-profiles). Don\u0026rsquo;t add weight_loss_pct, dev_time, or dtr_pct back to frontmatter — they\u0026rsquo;re computed in single.html and list.html from the four raw inputs. Don\u0026rsquo;t mix machine-specific guidance into methods.md. That page is meant to be universal across machines. Don\u0026rsquo;t modify the three completed roasts\u0026rsquo; measured data (green_weight_g, roasted_weight_g, time_to_fc, total_time, observations). The wrong-machine note in bean_notes is the documented record. ","permalink":"/roasts/claude/","summary":"Roasting section instructions This section logs coffee roasts on a Behmor 2000 AB Plus. The architecture separates universal roast science from machine-specific guidance so future machines can be added without template changes.\nFile layout content/roasts/00N-\u0026lt;origin\u0026gt;.md — one file per roast. Structured frontmatter holds plan + raw inputs + observations; the body is rendered from those fields, not from markdown prose. content/roasts/_TEMPLATE.md — scaffold for new roasts. Excluded from Tina collection via match.","title":"","type":"roasts"},{"content":"Cloud-Native | Full-Stack Dev | Developer \u0026amp; OSS Advocate 🥑\nResume PDF | Linktr.ee Summary Full-Stack engineer and Golang enthusiast with cloud-native, distributed systems expertise. Experienced in building scalable applications, event-driven microservices, and advocating for developer tools and frameworks. Flexible with a proven ability to collaborate globally across technical and non-technical teams. Self-motivated and empathetic, focusing on improving the developer experience and fostering open-source communities.\nCore Competencies Languages \u0026amp; Frameworks: Golang, Java, Spring Boot, Vue.js, React, TypeScript, Node.js Cloud \u0026amp; Infrastructure: Docker, Kubernetes, Helm, AWS, Google Cloud Event Streaming \u0026amp; Messaging: Kafka, gRPC, Protobuf, RabbitMQ Database Technologies: PostgreSQL, Redis, ElasticSearch, Google Cloud API DevOps \u0026amp; CI/CD: Git, Jenkins, CircleCI, Terraform, AWS CDK, Helm Developer Advocacy \u0026amp; Community: Technical Content Creation, Open-Source Maintenance, Public Speaking, Event Management Technical Leadership: Mentoring, People Management, Thought Leadership, Vision \u0026amp; Goal planning, Metric \u0026amp; KPI reporting Professional Experience Head of Developer Experience Aklivity | Mar 2023 - Nov 2024\nAklivity is an early-stage startup that enables event-driven solutions by proxying multiple API protocols onto Apache Kafka. I created and executed the vision for all aspects of community engagement and product documentation.\nLed the creation of a new, streamlined developer onboarding experience for the Zilla project , including documentation, examples repositories, and interactive demos. Enhanced community engagement through regular updates, feature announcements, and demo walkthroughs. Acted as the first point of contact for community troubleshooting and feedback, improving product quality based on user input. Represented the company at developer conferences and hackathons, creating product demos and handling booth responsibilities. OSS Maintainer Independent | Jan 2019 - Present\nGohlay : Creator of an open-source CLI tool written in Golang that interacts with Kafka to deliver scheduled messages. The tool is lightweight, scalable, and operates without external data dependencies. vue-pdf : Maintainer of a Vue.js library for rendering individual PDF pages as markup elements. awesome-data-engineering : Maintain the list, ensuring it remains up-to-date with the latest tools and technologies. Head of Developer Relations WSO2 | Jan 2022 - Mar 2023\nLed the Developer Relations team, building a thriving community around WSO2\u0026rsquo;s open-source and SaaS offerings, including the Ballerina programming language. Spearheaded the creation of the Ballerina Exercism.io track , improving language education and adoption. Coordinated global events, including hackathons and conferences. Spoke on technical topics at industry events. Collaborated with product teams and technical writers to enhance documentation and user experience across WSO2 products. Lead Solutions Engineer WSO2 | Feb 2021 - Jan 2022\nSupported sales teams in driving the adoption of WSO2 products by providing technical consultations and solutions. Secured the first customer for the Asgardeo SaaS platform . Authored an opinion article on identity management and tech debt . Senior Full Stack Developer BMW - Apps and Services | Feb 2020 - Feb 2021\nContributed to the development of BMW Charge Forward , a platform optimizing electric vehicle charging through real-time energy data. Built backend services using Node.js and NestJS, integrating with AWS services (Lambda, SQS, Kinesis) to handle streams of vehicle data. Senior Enterprise Application Engineer AFS Logistics | 2018 - Feb 2020\nLed development of a Transportation Management System using microservices architecture with Spring Boot, Golang, Kafka, Redis, PostgreSQL, and Kubernetes. Transportation Management System: A high-performance system with microservices deployed using Kubernetes and Helm. Key components included an A Rating System, a Document Manager, and a scalable Delayed Message Delivery service. Designed and implemented a Delayed Message Delivery Service that scaled efficiently using Kafka and Redis. Full Stack Mentor Thinkful Inc | 2018 - 2020\nMentored students in React.js and web design patterns, helping them overcome development hurdles and deepening their understanding of full-stack web development. Enterprise Application Engineer UberFreight | 2012 - 2018\nI maintained legacy applications, worked on improvements and new features of existing products like Real-time Search, and architected new applications like the Dynamic EDI Parser. Real-time Search: Added an ElasticSearch API on top of a monolithic Groovy on Grails, PostgreSQL, and Hibernate ORM application, reducing processing load and improving user-built query latency. Dynamic EDI Parser: Developed an algorithm to parse EDI files dynamically into object formats (XML, JSON) for easier integration and extensibility. Education B.S. in Computer Science (Cum Laude) Southern Wesleyan University\nAdditional Information Public Content: Highlights and links to my Blog, video, and documentation work Awards: Eagle Scout with Silver Palm Research: Path Planning using Dijkstra and Lightning Enhancement ","permalink":"/about/","summary":"Full-Stack Cloud Application Engineer | Developer Advocate | Passionate about OSS","title":"AJ Danelz","type":"page"}]