# HVAC Blueprint Analyzer — Session Handoff (2026-07-13, updated mid-day) Paste this whole file into a new chat to continue. It covers the 2026-07-11→13 session (GUI product-hardening + new detection layers) INCLUDING the 07-13 continuation (gold verdict, template sweep, Mac-review fixes). For deep engine/ledger history read the older `HANDOFF.md` first, then this. --- ## 0. GOLD REGRESSION — COMPLETE (all 4 buildings scored; DONE 07-13 11:11) The addenda gold regression (launched 07-13 08:04) finished all 4 buildings. Results in **`outputs/_gold_addenda/SUMMARY.md`** + `run.log`: | Building | Baseline | Addenda | Δ | FP | Mistag | Verdict | |---|---|---|---|---|---|---| | Lewis | 73.3 | **86.7** | **+13.4** | 28 | 0 | PASS | | Carroll | 90.4 | **85.5** | **−4.9** | 107 | 2 | **REGRESSION** | | 1326 | 70.4 | **81.7** | **+11.3** | 31 | 1 | PASS | | 23042 | 92.1 | **90.8** | **−1.3** | 52 | 13 | **REGRESSION** (marginal) | Aggregate recall across all 245 gold units: **~84.1% → ~86.1% (+2pts)** — but LOPSIDED: big double-digit wins on the two weaker buildings (Lewis, 1326), losses on the two already-strong ones (Carroll −4.9, 23042 −1.3). Carroll's FP also jumped (107) and 23042 mistag spiked (13 vs 0-2 elsewhere) — the addenda trade precision for recall unevenly. Decision below (default-OFF) stands and is reinforced by the completed 23042 row. **ACTION TAKEN (per the pre-agreed rule — any regression → revert):** `HVAC_PROMPT_ADDENDA` is now **default-OFF in code** (`analyze_blueprint_b.py::_prompt_addenda`, default flipped `"1"`→`"0"`). Set `HVAC_PROMPT_ADDENDA=1` to re-enable for experiments. Do NOT tune the addenda lines against gold — that poisons the gold set. ⚠ Open judgement call for the user: two double-digit WINS, two losses (−4.9, −1.3). Baselines are recorded numbers, not matched same-run controls — a $3 control run of Carroll with `HVAC_PROMPT_ADDENDA=0` (delete `pred_carroll/` first, then re-run the driver) would tell whether Carroll's 90.4 baseline was itself a lucky run vs. the addenda genuinely hurting it. User decides; default stays OFF until then. Note: the template sweep (§E below) now covers back-to-back twins deterministically — the main miss mode the addenda targeted, so the addenda may simply be redundant with a cleaner mechanism now. Regression run is COMPLETE — no resume needed. To re-run any building for a control, delete its `outputs/_gold_addenda/pred_/` and run: ```powershell $env:PYTHONIOENCODING='utf-8' .\.venv\Scripts\python.exe analyzer_b\hvac_pipeline\run_addenda_regression.py ``` (a building whose `pred_/` holds detections is skipped). --- ## 1. WHAT THIS SESSION SHIPPED (all uncommitted, all on branch `session/2026-07-08-thinkfix-validation-withers`) This session was driven by a repeating **internal focus-group loop** (persona panel critiques the live GUI on a fresh blueprint → fix → re-run until only positive). Panel + session reports live in **`focus-group-panel/`**. The test blueprint was **`raw_pdfs/Mac 100% CD.pdf`** (29 pages, never analyzed before); its results are cached in `outputs/Mac 100% CD/`. ### A. Two new revertible layers (built earlier this session) - **Runtime wargame layer** — `analyzer_b/hvac_pipeline/wargame_check.py` (REPORT-ONLY; per-page green/amber/red confidence + gold-free invariant checks). Runs in `run_qa`; `HVAC_WARGAME=0` disables. Flags fold into the GUI Flags tab. Tests: `test_wargame_check.py` (21 checks pass). - **Autopilot** — `wargame_autopilot.py` (OPT-IN, `HVAC_AUTOPILOT=1`, default OFF). Certifies green pages, tier-based counting, OCR gap-rescue. Snapshots review CSVs to `_pre_autopilot/`; undo = `--restore`. Tests: `test_wargame_autopilot.py` (30 checks pass). Validated it reproduced the user's 5 Withers review (28/28 + 7 junk) with zero human input. **Not yet run on a live GUI detection** — user hasn't flipped it on. ### B. GUI trust + usability fixes (focus-group rounds 1–3) All in `modules/pipeline.py`, `server.py`, `public/script.js` (now **v24**), `public/index.html`. Server run: `python server.py` (SYSTEM python), port 8000. - Truthful counts: summary leads **Verified units** + "to review / rejected / equipment" breakdown; gray noise never counted; the false "0 units found" banner is gated on the parsed total actually being 0. - **utf-8 console hardening** in server.py — a cp1252 `print()` of the engine's `→`/`×` used to CRASH a paid run. Fixed at startup. - **Content-hash cache reuse** (`_pdf_sha256.txt` + `_may_reuse_cache`): an identical re-upload reuses detections FREE (crash recovery no longer re-pays). `/api/recent` + `/api/reopen` + landing "Recent buildings" list reattach results after a restart. **Gotcha**: index.html is browser-cached; it's now served `no-cache`, but during dev use `?nocache=N` since `?v=` only busts script.js. - **Honest progress bar** — percent advances with tiles/snap/verify per page. - **Schedule intelligence**: missing-schedule flag names the exact unscheduled tags AND the pages that look like schedule sheets; adding those regions + "Read schedules" post-run updates capacities with ZERO re-detection (`/api/refresh-result`). Schedule-page hints shown at upload (`page_titles.py`, free text-layer scan). - **Auto floor names** from the PDF text layer (`page_titles.py`): picked pages become "Cellar"/"3rd Floor" not "Page 2". ⚠ This exposed + fixed a double-count bug: renaming orphaned old-name artifacts; ALL floor-file readers now dedupe by page number (newest mtime wins) via `_floor_files_newest()`. - **Product mode** (`HVAC_PRODUCT_MODE=1`): hides the dev-only Symbol Bank tab. - Junk-tag hygiene: word-fragment tags ("AC-TION" from "AIR CONDI-TIONING") arrive UNTICKED with a warning; real bare tags (AH-1 qty 4) untouched. - Advanced recipe knobs collapsed behind a `
`; model labels reworded to "Best accuracy (recommended)" / "Budget". - Mitsubishi P-series capacity codes (`PKFY-P08…` → 8 kBTU) added to `capacity.py`. ### C. Latest user-feedback batch (this is the freshest work) - **Processing time** on cached/reopened results shows the ORIGINAL run's wall-clock (from `timing_b.json`), never 0s. - **Type / Model / Location columns** done right in the Floors tab: `type_from_model()` resolves the real unit type from the model number via **`modules/unit_types_ref.json`** (53 families built from `Mitsubishi_Daikin_Indoor_HVAC_Units_v1.xlsx`; tolerates the plan's PFKY↔PKFY letter swap). Schedule text that is really a LOCATION ("APARTMENTS") moves to the Location column; new **Model** column added. - **ERV / MUA / EF / CU… = scope "equipment"**: shown muted with a badge, NEVER counted as indoor units (`equipment_units` in the summary). This is the ledger convention ("real equipment but not indoor"). - **"🗑 Delete yellows" button** beside Delete noise (confirm dialog, since some buildings hide real units in yellow). - **Rotatable Bank crop box** (from the prior message): the Symbol Bank green box rotates (`[`/`]` keys, ⟲/⟳ buttons, Shift=15°); on approve the crop is **de-rotated** so a diagonal unit body saves upright. Backend: `symbol_bank._crop_and_save` rotate-then-crop, proven with a synthetic 20°-stripe test. Angle recorded in `bank_decisions.json` + `crop_angle` in `symbol_library.json`. User confirmed working live 07-13. ### C2. 07-13 continuation fixes (while user reviewed Mac; script.js now v24) - **Floors & Units tab shows REAL floor names** ("Cellar", "3rd Floor") not the page ordinal; multi-floor "(sheet pN: …)" provenance moved to a hover tooltip (`populateFloorsTable`). Same tag on multiple floors = separate rows per floor, counted independently (verified live: Mac AH-1 on both). - **Add-unit button reset bug** (real bug): after a successful add the button stayed "✏️ drawing…" though draw-mode was off — a second add silently did nothing. Now resets label+highlight. NOTE: user's "relabel/delete missing, added box disappears" report was otherwise a STALE CACHED script.js — the features work; hard-refresh (Ctrl+F5) after every deploy, and remember the add-unit tag PROMPT discards the box on cancel/empty. - **PVFY family** added to `modules/unit_types_ref.json` (Multi-position Air Handler, CITY MULTI Ducted) — AH-5's Type no longer "-"; letter-swap tolerant (PFVY reads as PVFY); verified via `type_from_model`. ### D. GAP-HUNT — the new detection-recall layer (validated, shipped) `analyzer_b/hvac_pipeline/gap_hunt.py` productizes the user's "give Opus the image with all found boxes drawn transparent and ask what was missed" check. - Fable sees the plan with **only VERIFIED (green) boxes** drawn transparent green (bake-off A1: showing yellow/gray guesses BLINDS the hunter). Proposes missed indoor units → appended as **YELLOW review boxes** (additive only, never deletes) + `gap_hunt_flags.csv`. - **2×2 tiling** for big sheets (whole-page scored 0/3 on a seeded test — units too small at the hi-res tier). Proposals then **OCR-snapped** to printed tags (Track-C move) and deduped against existing greens. - Runs in `run_qa` before wargame. `HVAC_GAP_HUNT=0` disables. Needs the `anthropic` sdk (now installed in the venv) + `ANTHROPIC_API_KEY` (present). - **Seeded validation** (removed 3 verified units from Mac p5, made it re-find): whole-page **0/3** → tiled **1/3** → tiled+snap **2/3 at 14-16px precision**, 3 extra yellows/floor. Cost ~$0.13/floor (~21k in / 12k out tokens, 4 tiles). - Also this batch: **prompt addenda** (`HVAC_PROMPT_ADDENDA`, §0 — now default-OFF after the gold verdict) — sharpened back-to-back-wall rule + A3 "orient first" line in `analyze_blueprint_b.py::_prompt_addenda()`; and **room_coverage_check.py** now emits neigh≥3 edge-hole candidates (model-confirmed only) to catch back-to-back/corner misses. ### E. TEMPLATE SWEEP — deterministic recall net (built+validated 07-13 PM) `analyzer_b/hvac_pipeline/template_sweep.py` — born from the user's Mac 3rd- floor review: she hand-added 6 units the model missed; ALL were back-to-back twins / wall-flush, all the SAME CAD glyph as the 67 found units. So: the run's own VERIFIED boxes become OpenCV templates (body-only via `body_crop`, tag text stripped), swept at 4 rotations at half-res, NMS, then the gate that makes it precise: **a proposal is kept only if an UNCLAIMED printed tag word (not inside any live review box) sits within 260px**. Grille/diffuser glyphs (matched 0.76–0.81 — indistinguishable from units by score!) die because their nearest tag is claimed; real misses live because every one has its own tag beside it. OCR is **escalating** (the proven snap trick): whole-page pass reads only ~half the underlined tags → targeted 2× crops around top-40 hits with underline removal, psm 11 + psm 6 (11 splits "AH-2" into pieces) + rot90 pass; em/en-dashes normalized ("AH—1"); families must be ≥2 letters (kills 'Ø 4"'→"O-4" misreads). ADDITIVE-only yellows (reason=template_sweep) + `template_sweep_flags.csv` → GUI Flags tab. `HVAC_TEMPLATE_SWEEP=0` disables. Runs in run_qa BEFORE gap_hunt (hunter's dedupe then skips sweep-queued yellows). ~2-3 min/page, $0, needs tesseract. - **Validation (on COPIES, live Mac CSVs untouched)**: pre-review p5 state → recovers 4/6 of her misses (both members of the both-missed pair dead-on, the 90°-vertical one, one ~120px-offset) + **2 real CH-2 cabinet heaters nobody found**, 0 junk. The 2 unrecovered have OCR-illegible tags (one overprinted by room text) — gap-hunt's territory; thresholds NOT tuned to chase n=2. Post-review state (both pages) → **5/5 proposals real** (CH-2 ×5), 100% precision. Tests: `test_template_sweep.py`, 21 checks pass. - The 5 CH-2 proposals live only in scratch copies (`C:\hvac_work\ sweep_validate`) — the sweep has NOT been run on the live Mac out-dir yet (user's call, after she finishes reviewing). - Fixed in passing: run_qa's `--render-pdf` was appended to `jobs[-1]` (whatever QA job happened to be last → argparse crash → silent layer skip on crisp-PDF runs); now targeted explicitly to the scripts that accept it. template_sweep also accepts `--render-pdf` (coords live in its raster space). - CH + EUH added to `_EQUIPMENT_FAMILIES` (cabinet/electric unit heaters show muted as equipment, never counted as indoor). --- ## 2. IMMEDIATE QUEUE FOR THE NEW CHAT (in order) 1. **Finish the gold regression** (§0): check 23042's row in `outputs/_gold_addenda/SUMMARY.md` (was in flight ~10:03). The revert (default-OFF) is already DONE; 23042's row just completes the record. Then user decides on the optional $3 Carroll control run (§0 caveat). 2. **Server restart pending**: `modules/pipeline.py` changed (template sweep in run_qa, CH/EUH equipment, PVFY ref) — the running server doesn't have them. Restart drops in-memory JOBS; "Recent buildings" reattaches free. Don't restart while she has unsaved floor edits open. 3. **Mac review in progress → first-contact ledger row 7** (`analyzer_b/results/first_contact_ledger.csv`). 3rd Floor SAVED (67 confirmed + 6 user-added, all 6 = back-to-back twins the model missed — the evidence that drove §E). Cellar state unknown. Her 6 added boxes + the CH-2s are prime bank/corrections_log harvest once done. 4. **Offer on the table**: run template_sweep on the LIVE Mac out-dir so the 5 real CH-2 yellows appear in her GUI (additive-only; after her review): ```powershell .\.venv\Scripts\python.exe analyzer_b\hvac_pipeline\template_sweep.py ` --out-dir "outputs\Mac 100% CD" --pdf "uploads\.pdf" ``` 5. **Git**: everything is uncommitted on `session/2026-07-08-thinkfix-validation-withers`. User decides when to commit. New untracked files to include: `wargame_check.py`, `wargame_autopilot.py`, `gap_hunt.py`, `template_sweep.py`, `test_template_sweep.py`, `page_titles.py`, `run_addenda_regression.py`, `test_wargame_*.py`, `modules/unit_types_ref.json`, `wargames_hvac.md`, `focus-group-panel/`. Before `git add`: scan staged for `.env|pdf|png|zip|pt`; NEVER add the corrupted `HVAC-Blueprint-Analyzer*` root junk (U+F05C filenames, has a stray live-key `.env`). ## 3. OPEN ROADMAP ASKS (proposed, NOT built — need user go-ahead) - Category expansion beyond indoor units: the "What to look for" dropdown (diffusers, electric heaters, etc.) is still a **placeholder — only Indoor units works**. Each new category needs its own prompt/verify/gold-building loop (the separate `C:\hvac_fable` rebuild targets 11 categories properly). - Enterprise (from harsh-panel round 3): multi-user/login; Excel export with per-box coordinates+provenance; a first-run onboarding tour. - Track C (Fable-as-proposer, highest upside): the full version of gap-hunt's snap idea — Fable proposes ALL units, OCR grounds them. `fable_propose.py` exists; probe was 11/11 & 6/6 on the floors that broke Gemini. - Template-sweep v2 ideas (NOT built; don't tune on n=1): cross-floor template seeding; a "strong-hit (≥0.85) + claimed-tag-nearby" rescue tier for the two OCR-illegible Mac misses; box-recenter onto the unit body when the kept hit is the tag-adjacent one (~120px offset case). ## 4. ENVIRONMENT (read before running) - Windows 11, PowerShell (no `&&`; write scratchpad `.py` for multi-line `python -c`; engine output through Select-Object can exit 255 benignly → redirect to a log). Bash tool is BROKEN this session. - Engine python: `.venv\Scripts\python.exe` (has cv2/fitz/genai/anthropic). GUI server: SYSTEM `python server.py` (has fastapi). Split is intentional. - After editing `public/*`: bump `script.js?v=N` in index.html + hard-refresh. After editing `server.py`/`modules/*`: restart the server. - `.env` has GEMINI + ANTHROPIC (both present/working) + OPENROUTER (small cap). - Heavy IO → `C:\hvac_work` (OneDrive Errno 22). - Feature-flag summary (all default to current behavior): `HVAC_WARGAME=0` off wargame · `HVAC_AUTOPILOT=1` on autopilot (default off) · `HVAC_GAP_HUNT=0` off gap-hunt · `HVAC_TEMPLATE_SWEEP=0` off template sweep (default ON, $0) · `HVAC_PROMPT_ADDENDA=1` re-enable addenda (**default OFF since the 07-13 gold verdict**) · `HVAC_PRODUCT_MODE=1` hide Bank tab. ## 5. KEY FILES ADDED/CHANGED THIS SESSION - NEW: `analyzer_b/hvac_pipeline/{wargame_check,wargame_autopilot,gap_hunt, template_sweep,test_template_sweep,page_titles,run_addenda_regression, test_wargame_check,test_wargame_autopilot}.py` - NEW: `modules/unit_types_ref.json` (+PVFY 07-13), `focus-group-panel/`, `wargames_hvac.md`, this file. - CHANGED: `modules/pipeline.py` (type/scope/model resolver, cache-hash, floor-name fill, dedupe, refresh-result, template-sweep+gap-hunt+wargame jobs, --render-pdf targeting fix, equipment counting incl. CH/EUH), `modules/capacity.py` (P-series), `modules/symbol_bank.py` (rotated crop), `server.py` (utf-8, /api/recent|reopen|refresh-result|config, no-cache root), `public/script.js` (v24: floor names in Floors&Units, add-unit button reset), `public/index.html`, `analyzer_b/analyze_blueprint_b.py` (`_prompt_addenda`, default OFF 07-13), `analyzer_b/hvac_pipeline/room_coverage_check.py` (neigh≥3). - Scratch validation artifacts (not for git): `C:\hvac_work\sweep_validate\`, `C:\hvac_work\sweep_recall\`.