Official implementation of:
"面向外观专利审查的专利级多视图检索方法"
PatMV is a patent-level multi-view retrieval framework designed for U.S. design patent novelty examination research.
Unlike traditional design patent retrieval methods that treat individual patent views independently, PatMV models a complete design patent as a multi-view image set and learns patent-level visual representations by explicitly exploiting cross-view geometric relationships.
A design patent usually contains multiple views (e.g., front, rear, side, top, bottom, and perspective views), which jointly describe the three-dimensional appearance of an object.
Existing image-level retrieval paradigms may lose complementary information among different views. To address this issue, PatMV introduces a patent-level multi-view representation learning framework.
The main components include:
-
Patent-level visual encoding
- Treats all views of one patent as a unified input unit.
- Learns a global representation for each complete patent.
-
View-aware feature modeling
- Introduces learnable view identifier embeddings together with spatial position embeddings.
- Preserves both spatial information and view-source information.
-
Global Patch-level Cross-View Attention
- Inserts cross-view attention modules into the Vision Transformer architecture.
- Enables fine-grained interactions among image patches from different patent views.
-
Gated cross-view feature fusion
- Uses zero-initialized learnable gates to smoothly integrate cross-view information while maintaining pretrained feature stability.
-
Multi-objective contrastive learning
- Jointly optimizes:
- patent-text alignment;
- patent-single-view consistency.
- Jointly optimizes:
The PatMV dataset is publicly available at:
Hugging Face Dataset
https://huggingface.co/datasets/blueingman/PatMV
The dataset is constructed from official USPTO examination records and is designed for research on U.S. design patent novelty analysis.
It contains:
- examiner-confirmed similar design patent pairs;
- multi-view patent images;
- image-level captions;
- patent-level captions;
- distractor patents for retrieval evaluation.
Please refer to the dataset card for details.
PatMV/
│
├── model/
│ └── Model architectures
│
├── dataset/
│ └── Dataset loading and preprocessing
│
├── loss/
│ └── Contrastive learning objectives
│
├── train/
│ └── Training scripts
│
├── validate/
│ └── Validation and evaluation scripts
│
├── inference/
│ └── Retrieval inference
│
└── data_processing/
└── Dataset conversion and preparation scripts
The released dataset provides processed metadata and images.
After downloading the dataset from Hugging Face, organize files as:
data/
├── images/
├── query_metadata.json
└── gallery_metadata.json
The scripts in data_processing/ provide utilities for converting and preparing patent retrieval data.
To train PatMV:
python train.pyTraining configuration can be modified according to your environment.
The training procedure optimizes a joint contrastive objective consisting of:
- Patent-level visual-text alignment;
- Patent-level and single-view visual consistency.
To evaluate retrieval performance:
python inference.pyThe evaluation pipeline supports patent-level retrieval under gallery distractors.
Common retrieval metrics include:
- Hit@K
- mAP
The pretrained checkpoint will be released separately.
If you use this project or the PatMV dataset in your research, please cite:
Wang Yi-wei et al.
"面向外观专利审查的专利级多视图检索方法."
Accepted by CCIR 2026.
This repository is released for academic research purposes only.
Users should ensure compliance with applicable intellectual property regulations and dataset usage policies.
This work builds upon previous advances in:
- Vision Transformer;
- Contrastive Language-Image Pretraining;
- Patent-oriented multimodal representation learning.
