T

gbanyanandClaude Opus 4.7 0471e36fd4 Paper A v3.16: remove unsupported visual-inspection / sanity-sample claims

User review of the v3.15 Sanity Sample subsection revealed that the
paper's claim of "inter-rater agreement with the classifier in all 30
cases" (Results IV-G.4) was not backed by any data artifact in the
repository. Script 19 exports a 30-signature stratified sample to
reports/pixel_validation/sanity_sample.csv, but that CSV contains
only classifier output fields (stratum, sig_id, cosine, dhash_indep,
pixel_identical, closest_match) and no human-annotation column, and
no subsequent script computes any human--classifier agreement metric.
User confirmed that the only human annotation in the project was
the YOLO training-set bounding-box labeling; signature classification
(stamped vs hand-signed) was done entirely by automated numerical
methods. The 30/30 sanity-sample claim was therefore factually
unsupported and has been removed.

Investigation additionally revealed that the "independent visual
inspection of randomly sampled Firm A reports reveals pixel-identical
signature images...for many of the sampled partners" framing used as
the first strand of Firm A's replication-dominated evidence (Section
III-H first strand, Section V-C first strand, and the Conclusion
fourth contribution) had the same provenance problem: no human
visual inspection was performed. The underlying FACT (that Firm A
contains many byte-identical same-CPA signature pairs) is correct
and fully supported by automated byte-level pair analysis (Script 19),
but the "visual inspection" phrasing misrepresents the provenance.

Changes:

1. Results IV-G.4 "Sanity Sample" subsection deleted entirely
(results_v3.md L271-273).

2. Methodology III-K penultimate paragraph describing the 30-signature
manual visual sanity inspection deleted (methodology_v3.md L259).

3. Methodology Section III-H first strand (L152) rewritten from
"independent visual inspection of randomly sampled Firm A reports
reveals pixel-identical signature images...for many of the sampled
partners" to "automated byte-level pair analysis (Section IV-G.1)
identifies 145 Firm A signatures that are byte-identical to at
least one other same-CPA signature from a different audit report,
distributed across 50 distinct Firm A partners (of 180 registered); 35 of these byte-identical matches span different fiscal years."
All four numbers verified directly from the signature_analysis.db
database via pixel_identical_to_closest = 1 filter joined to
accountants.firm.

4. Discussion V-C first strand (L41) rewritten analogously to refer
to byte-level pair evidence with the same four verified numbers.

5. Conclusion fourth contribution (L21) rewritten to "byte-level
pair analysis finding of 145 pixel-identical calibration-firm
signatures across 50 distinct partners (Section IV-G.1)."

6. Abstract (L5): "visual inspection and accountant-level mixture
evidence..." rewritten as "byte-level pixel-identity evidence
(145 signatures across 50 partners) and accountant-level mixture
evidence..." Abstract now at 250/250 words.

7. Introduction (L55): "visual-inspection evidence" relabeled
"byte-level pixel-identity evidence" for internal consistency.

8. Methodology III-H penultimate (L164): "validation role is played
by the visual inspection" relabeled "validation role is played
by the byte-level pixel-identity evidence" for consistency.

All substantive claims are preserved and now back-traceable to
Script 19 output and the signature_analysis.db pixel_identical_to_closest
flag. This correction brings the paper's descriptive language into
strict alignment with its actual methodology, which is fully
automated (except for YOLO training annotation, disclosed in
Methodology Section III-B).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

2026-04-25 01:14:13 +08:00

paper

Paper A v3.16: remove unsupported visual-inspection / sanity-sample claims

2026-04-25 01:14:13 +08:00

signature_analysis

Paper A v3.7: demote BD/McCrary to density-smoothness diagnostic; add Appendix A

2026-04-21 14:32:50 +08:00

signature-comparison

Complete PP-OCRv5 research and v4 vs v5 comparison

2025-11-27 11:21:55 +08:00

test_results

Complete PP-OCRv5 research and v4 vs v5 comparison

2025-11-27 11:21:55 +08:00

.gitignore

Add Deloitte distribution & independent dHash analysis scripts

2026-04-20 21:34:24 +08:00

check_rejected_for_missing.py

Add PaddleOCR masking and region detection pipeline

2025-10-28 22:28:18 +08:00

COMMIT_SUMMARY.md

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

CURRENT_STATUS.md

Complete OpenCV Method 3 implementation with 86.5% handwriting retention

2025-11-27 10:35:46 +08:00

extract_handwriting.py

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

extract_pages_from_csv.py

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

extract_signatures_hybrid.py

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

extract_signatures_paddleocr_improved.py

Complete OpenCV Method 3 implementation with 86.5% handwriting retention

2025-11-27 10:35:46 +08:00

extract_signatures_vlm.py

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

extract_signatures_yolo.py

Add Paper A (IEEE TAI) complete draft with Firm A-calibrated dual-method classification

2026-04-06 23:05:33 +08:00

HOW_TO_CONTINUE.txt

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

NEW_SESSION_HANDOFF.md

Complete OpenCV Method 3 implementation with 86.5% handwriting retention

2025-11-27 10:35:46 +08:00

NEW_SESSION_PROMPT.txt

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

paddleocr_client.py

Add PaddleOCR masking and region detection pipeline

2025-10-28 22:28:18 +08:00

paddleocr_server_v5.py

Complete OpenCV Method 3 implementation with 86.5% handwriting retention

2025-11-27 10:35:46 +08:00

PADDLEOCR_STATUS.md

Add PaddleOCR masking and region detection pipeline

2025-10-28 22:28:18 +08:00

PP_OCRV5_RESEARCH_FINDINGS.md

Complete PP-OCRv5 research and v4 vs v5 comparison

2025-11-27 11:21:55 +08:00

PROJECT_DOCUMENTATION.md

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

README_hybrid_extraction.md

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

README_page_extraction.md

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

README.md

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

SAM3_RESEARCH_FINDINGS.md

Add Paper A (IEEE TAI) complete draft with Firm A-calibrated dual-method classification

2026-04-06 23:05:33 +08:00

SESSION_CHECKLIST.md

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

SESSION_INIT.md

Add hybrid signature extraction with name-based verification

2025-10-26 23:39:52 +08:00

test_mask_and_detect.py

Add PaddleOCR masking and region detection pipeline

2025-10-28 22:28:18 +08:00

test_opencv_advanced.py

Complete OpenCV Method 3 implementation with 86.5% handwriting retention

2025-11-27 10:35:46 +08:00

test_opencv_separation.py

Complete OpenCV Method 3 implementation with 86.5% handwriting retention

2025-11-27 10:35:46 +08:00

test_paddleocr_client.py

Add PaddleOCR masking and region detection pipeline

2025-10-28 22:28:18 +08:00

test_paddleocr.py

Add PaddleOCR masking and region detection pipeline

2025-10-28 22:28:18 +08:00

test_pp_ocrv5_api.py

Complete PP-OCRv5 research and v4 vs v5 comparison

2025-11-27 11:21:55 +08:00

test_v4_full_pipeline.py

Complete PP-OCRv5 research and v4 vs v5 comparison

2025-11-27 11:21:55 +08:00

test_v5_full_pipeline.py

Complete PP-OCRv5 research and v4 vs v5 comparison

2025-11-27 11:21:55 +08:00

visualize_v5_results.py

Complete PP-OCRv5 research and v4 vs v5 comparison

2025-11-27 11:21:55 +08:00

yolo_extract_from_index.py

Add Paper A (IEEE TAI) complete draft with Firm A-calibrated dual-method classification

2026-04-06 23:05:33 +08:00

yolo_full_scan.py

Add Paper A (IEEE TAI) complete draft with Firm A-calibrated dual-method classification

2026-04-06 23:05:33 +08:00

README.md

PDF Signature Extraction System

Automated extraction of handwritten Chinese signatures from PDF documents using hybrid VLM + Computer Vision approach.

Quick Start

Step 1: Extract Pages from CSV

cd /Volumes/NV2/pdf_recognize
source venv/bin/activate
python extract_pages_from_csv.py

Step 2: Extract Signatures

python extract_signatures_hybrid.py

Documentation

PROJECT_DOCUMENTATION.md - Complete project history, all approaches tested, detailed results
README_page_extraction.md - Page extraction documentation
README_hybrid_extraction.md - Hybrid signature extraction documentation

Current Performance

Test Dataset: 5 PDF pages

Signatures expected: 10
Signatures found: 7
Precision: 100% (no false positives)
Recall: 70%

Key Features

✅ Hybrid Approach: VLM name extraction + CV detection + VLM verification ✅ Name-Based: Signatures saved as signature_周寶蓮.png ✅ No False Positives: Name-specific verification filters out dates, text, stamps ✅ Duplicate Prevention: Only one signature per person ✅ Handles Both: PDFs with/without text layer

File Structure

extract_pages_from_csv.py          # Step 1: Extract pages
extract_signatures_hybrid.py       # Step 2: Extract signatures (CURRENT)
README.md                          # This file
PROJECT_DOCUMENTATION.md           # Complete documentation
README_page_extraction.md          # Page extraction guide
README_hybrid_extraction.md        # Signature extraction guide

Requirements

Python 3.9+
PyMuPDF, OpenCV, NumPy, Requests
Ollama with qwen2.5vl:32b model
Ollama instance: http://192.168.30.36:11434

Data

Input: /Volumes/NV2/PDF-Processing/master_signatures.csv (86,073 rows)
PDFs: /Volumes/NV2/PDF-Processing/total-pdf/batch_*/
Output: /Volumes/NV2/PDF-Processing/signature-image-output/

Status

✅ Page extraction: Tested with 100 files, working ✅ Signature extraction: Tested with 5 files, 70% recall, 100% precision ⏳ Large-scale testing: Pending ⏳ Full dataset (86K files): Pending

See PROJECT_DOCUMENTATION.md for complete details.