The program extracts data from measurement strips contained within a PDF and converts it into an Excel file. The process involves first converting the PDF pages into images, then using OCR (Optical Character Recognition) through Pytesseract to extract the data.
DocuScan allows users to process measurement equipment data with associated measuring dates directly from a PDF. The program includes options to set DPI for the PDF-to-image conversion (recommended to match the DPI of the scanner or printer for accuracy). After selecting the desired PDF and pressing "Convert," the program scans all measurement strips within the document. Each page from the PDF is converted into an individual sheet in an Excel file.
To get a .exe follow these steps:
-
Clone the repository:
git clone https://github.com/Scurbs/DocuScan.git
-
Install dependencies:
For good practice it is recommended to use a virtual environment
pip install -r requirements.txt -
Install Tesseract for Windows:
Download Tesseract from UB-Mannheim for Windows Place the folder "Tesseract" into the same folder where main.py is.
-
Edit DocuScan.spec: If you want to use the debug mode fully, then you have to set the flag on true for console
-
Run the main.spec file for building the .exe
pyinstaller DocuScan.spec -
Place the .exe in the same folder where main.py is. You can now use the program !
The Start page offers four main options:
- Select File: Choose a PDF file to process.
- Convert: Begin converting the selected PDF to an Excel file.
- Settings: Adjust conversion and debugging settings.
Under Settings, set the DPI (recommended to match the DPI of your scanner or printer) for accurate conversion of PDF pages to images. You can also enable Debug Mode, which opens a console to display debug statements during the conversion process
Ensure a file is selected before starting the conversion process. Once selected, press "Convert" to begin transforming the PDF into an Excel file.
The generated Excel file includes:
- Timestamp: From each measurement there is the according timestamp
- Particle size: For each measuring stripe there a two colums for the particle size 0.5µm and 0.3µm
- Serial Number: The serial number for each measuring stripe is also displayed
- The mean values for each particle size
- Checks if the value is within a range of 15%
This project is licensed under the MIT License - see the LICENSE file for details.
This project utilizes the following third-party tools:
- pytesseract: A Python wrapper for Tesseract OCR, licensed under the Apache License 2.0.
- Tesseract OCR: An OCR engine developed by Google, also licensed under the Apache License 2.0.
Please note that by using this project, you must also comply with the terms of the Apache License 2.0 for pytesseract and Tesseract OCR. A copy of the Apache License is available here.
Created by Janes Pozar & Avdo Muminovic. For additional questions, please reach out through janes.pozar@students.fhnw.ch or visit the project repository.