T

gbanyanandClaude Opus 4.7 723a3f6eaf Rewrite §III v7: anchor-based ICCR framework + composition-decomp finding

Major §III restructuring after codex rounds 29-34 demolished the
distributional path to thresholds (Scripts 39b-39e prove (cos, dHash)
multimodality is composition-driven + integer-tie artefact).

v4.0 pivots to anchor-based multi-level inter-CPA coincidence-rate
(ICCR) calibration via Scripts 40b, 43, 44, 45, 46:

- §III-G: scope justification rewritten (LOOO + Firm A case study +
  within-firm collision structure; dropped "smallest scope rejects
  unimodality" rationale); added sample-size reconciliation
  (150,442 descriptor-complete vs 150,453 vector-complete; 437
  accountant-level vs 468 all)
- §III-I: new sub-section I.4 composition decomposition (2x2 factorial
  centred + jittered Big-4 pooled dh p=0.35); I.5 conclusion of no
  natural threshold
- §III-J: K=3 recast as firm-compositional descriptive partition
  (not three mechanism clusters); bridge to §III-L.4 cross-firm
  hit matrix added
- §III-K: Score 1 reframed as firm-composition position score
- §III-L: NEW major sub-section — anchor-based threshold calibration
  with L.0 methodology, L.1 per-comparison ICCR (replicates v3
  cos>0.95 -> 0.0006; new dh<=5 -> 0.0013; joint -> 0.00014),
  L.2 pool-normalised per-signature ICCR (any-pair HC 11.02%;
  per-firm A 25.94% vs B/C/D <1.5%), L.3 doc-level ICCR (HC 18%;
  HC+MC 34%), L.4 firm heterogeneity logistic OR 0.01-0.05 +
  cross-firm hit matrix (98-100% within-firm), L.5 alert-rate
  sensitivity (HC threshold locally sensitive not plateau-stable),
  L.6 observed deployed alert rate excess over inter-CPA proxy
- §III-M: NEW sub-section — multi-tool validation strategy under
  unsupervised setting; 9 partial-evidence diagnostics each with
  disclosed untested assumption; positioning as anchor-calibrated
  screening framework with human-in-the-loop review, NOT validated
  forensic detector
- Terminology: "FAR" replaced with "inter-CPA coincidence rate
  (ICCR)" throughout; primary metric name change documented in
  §III-L.0
- Provenance table: ~35 new rows for Scripts 39b-e/40b/43-46;
  "key numerical claims" instead of "every numerical claim"
- Removed v2-v6 internal changelog metadata; v7 draft note added

Codex round-32 SOUND_WITH_QUALIFICATIONS, round-33 GO_WITH_REVISIONS,
round-34 READY_WITH_NARROW_FIXES (all 8 patches applied).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

2026-05-13 17:27:01 +08:00

.planning

Update STATE.md: Phase 1 complete, Phase 2 awaiting user review

2026-05-12 15:24:03 +08:00

paper

Rewrite §III v7: anchor-based ICCR framework + composition-decomp finding

2026-05-13 17:27:01 +08:00

signature_analysis

Add Script 46: alert-rate sensitivity / threshold-plateau analysis

2026-05-13 16:46:08 +08:00

signature-comparison

Complete PP-OCRv5 research and v4 vs v5 comparison

2025-11-27 11:21:55 +08:00

test_results

Complete PP-OCRv5 research and v4 vs v5 comparison

2025-11-27 11:21:55 +08:00

.gitignore

Add Deloitte distribution & independent dHash analysis scripts

2026-04-20 21:34:24 +08:00

check_rejected_for_missing.py

Add PaddleOCR masking and region detection pipeline

2025-10-28 22:28:18 +08:00

COMMIT_SUMMARY.md

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

CURRENT_STATUS.md

Complete OpenCV Method 3 implementation with 86.5% handwriting retention

2025-11-27 10:35:46 +08:00

extract_handwriting.py

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

extract_pages_from_csv.py

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

extract_signatures_hybrid.py

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

extract_signatures_paddleocr_improved.py

Complete OpenCV Method 3 implementation with 86.5% handwriting retention

2025-11-27 10:35:46 +08:00

extract_signatures_vlm.py

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

extract_signatures_yolo.py

Add Paper A (IEEE TAI) complete draft with Firm A-calibrated dual-method classification

2026-04-06 23:05:33 +08:00

HOW_TO_CONTINUE.txt

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

NEW_SESSION_HANDOFF.md

Complete OpenCV Method 3 implementation with 86.5% handwriting retention

2025-11-27 10:35:46 +08:00

NEW_SESSION_PROMPT.txt

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

paddleocr_client.py

Add PaddleOCR masking and region detection pipeline

2025-10-28 22:28:18 +08:00

paddleocr_server_v5.py

Complete OpenCV Method 3 implementation with 86.5% handwriting retention

2025-11-27 10:35:46 +08:00

PADDLEOCR_STATUS.md

Add PaddleOCR masking and region detection pipeline

2025-10-28 22:28:18 +08:00

PP_OCRV5_RESEARCH_FINDINGS.md

Complete PP-OCRv5 research and v4 vs v5 comparison

2025-11-27 11:21:55 +08:00

PROJECT_DOCUMENTATION.md

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

README_hybrid_extraction.md

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

README_page_extraction.md

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

README.md

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

SAM3_RESEARCH_FINDINGS.md

Add Paper A (IEEE TAI) complete draft with Firm A-calibrated dual-method classification

2026-04-06 23:05:33 +08:00

SESSION_CHECKLIST.md

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

SESSION_INIT.md

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

test_mask_and_detect.py

Add PaddleOCR masking and region detection pipeline

2025-10-28 22:28:18 +08:00

test_opencv_advanced.py

Complete OpenCV Method 3 implementation with 86.5% handwriting retention

2025-11-27 10:35:46 +08:00

test_opencv_separation.py

Complete OpenCV Method 3 implementation with 86.5% handwriting retention

2025-11-27 10:35:46 +08:00

test_paddleocr_client.py

Add PaddleOCR masking and region detection pipeline

2025-10-28 22:28:18 +08:00

test_paddleocr.py

Add PaddleOCR masking and region detection pipeline

2025-10-28 22:28:18 +08:00

test_pp_ocrv5_api.py

Complete PP-OCRv5 research and v4 vs v5 comparison

2025-11-27 11:21:55 +08:00

test_v4_full_pipeline.py

Complete PP-OCRv5 research and v4 vs v5 comparison

2025-11-27 11:21:55 +08:00

test_v5_full_pipeline.py

Complete PP-OCRv5 research and v4 vs v5 comparison

2025-11-27 11:21:55 +08:00

visualize_v5_results.py

Complete PP-OCRv5 research and v4 vs v5 comparison

2025-11-27 11:21:55 +08:00

yolo_extract_from_index.py

Add Paper A (IEEE TAI) complete draft with Firm A-calibrated dual-method classification

2026-04-06 23:05:33 +08:00

yolo_full_scan.py

Add Paper A (IEEE TAI) complete draft with Firm A-calibrated dual-method classification

2026-04-06 23:05:33 +08:00

README.md

PDF Signature Extraction System

Automated extraction of handwritten Chinese signatures from PDF documents using hybrid VLM + Computer Vision approach.

Quick Start

Step 1: Extract Pages from CSV

cd /Volumes/NV2/pdf_recognize
source venv/bin/activate
python extract_pages_from_csv.py

Step 2: Extract Signatures

python extract_signatures_hybrid.py

Documentation

PROJECT_DOCUMENTATION.md - Complete project history, all approaches tested, detailed results
README_page_extraction.md - Page extraction documentation
README_hybrid_extraction.md - Hybrid signature extraction documentation

Current Performance

Test Dataset: 5 PDF pages

Signatures expected: 10
Signatures found: 7
Precision: 100% (no false positives)
Recall: 70%

Key Features

✅ Hybrid Approach: VLM name extraction + CV detection + VLM verification ✅ Name-Based: Signatures saved as signature_周寶蓮.png ✅ No False Positives: Name-specific verification filters out dates, text, stamps ✅ Duplicate Prevention: Only one signature per person ✅ Handles Both: PDFs with/without text layer

File Structure

extract_pages_from_csv.py          # Step 1: Extract pages
extract_signatures_hybrid.py       # Step 2: Extract signatures (CURRENT)
README.md                          # This file
PROJECT_DOCUMENTATION.md           # Complete documentation
README_page_extraction.md          # Page extraction guide
README_hybrid_extraction.md        # Signature extraction guide

Requirements

Python 3.9+
PyMuPDF, OpenCV, NumPy, Requests
Ollama with qwen2.5vl:32b model
Ollama instance: http://192.168.30.36:11434

Data

Input: /Volumes/NV2/PDF-Processing/master_signatures.csv (86,073 rows)
PDFs: /Volumes/NV2/PDF-Processing/total-pdf/batch_*/
Output: /Volumes/NV2/PDF-Processing/signature-image-output/

Status

✅ Page extraction: Tested with 100 files, working ✅ Signature extraction: Tested with 5 files, 70% recall, 100% precision ⏳ Large-scale testing: Pending ⏳ Full dataset (86K files): Pending

See PROJECT_DOCUMENTATION.md for complete details.