Showing content from https://github.com/PaddlePaddle/PaddleOCR/tree/release/2.5 below:
GitHub - PaddlePaddle/PaddleOCR at release/2.5
English | 简体中文
PaddleOCR aims to create multilingual, awesome, leading, and practical OCR tools that help users train better models and apply them into practice.
- 🔥2022.7 Release OCR scene application collection
- PaddleOCR scene application covers general, manufacturing, finance, transportation industry of the main OCR vertical applications, including digital tube, LCD screen character, license plate, high-precision SVTR model, etc. 7 vertical models.
- 🔥2022.5.9 Release PaddleOCR release/2.5
- Release PP-OCRv3: With comparable speed, the effect of Chinese scene is further improved by 5% compared with PP-OCRv2, the effect of English scene is improved by 11%, and the average recognition accuracy of 80 language multilingual models is improved by more than 5%.
- Release PPOCRLabelv2: Add the annotation function for table recognition task, key information extraction task and irregular text image.
- Release interactive e-book "Dive into OCR", covers the cutting-edge theory and code practice of OCR full stack technology.
- 2021.12.21 Release PaddleOCR release/2.4
- Release 1 text detection algorithm (PSENet), 3 text recognition algorithms (NRTR、SEED、SAR).
- Release 1 key information extraction algorithm SDMGR and 3 DocVQA algorithms (LayoutLM, LayoutLMv2, LayoutXLM).
- 2021.9.7 Release PaddleOCR release/2.3
- Release PP-OCRv2. The inference speed of PP-OCRv2 is 220% higher than that of PP-OCR server in CPU device. The F-score of PP-OCRv2 is 7% higher than that of PP-OCR mobile.
- 2021.8.3 Release PaddleOCR release/2.2
- Release a new structured documents analysis toolkit, i.e., PP-Structure, support layout analysis and table recognition (One-key to export chart images to Excel files).
- more
PaddleOCR support a variety of cutting-edge algorithms related to OCR, and developed industrial featured models/solution PP-OCR and PP-Structure on this basis, and get through the whole process of data production, model training, compression, inference and deployment.
It is recommended to start with the “quick start” in the document tutorial
-
For international developers, we regard PaddleOCR Discussions as our international community platform. All ideas and questions can be discussed here in English.
-
For Chinese develops, Scan the QR code below with your Wechat, you can join the official technical discussion group. For richer community content, please refer to 中文README, looking forward to your participation.
PP-OCR Series Model List(Update on September 8th)
PP-OCRv3 Chinese model PP-OCRv3 English model PP-OCRv3 Multilingual model PP-Structure
- layout analysis + table recognition
- SER (Semantic entity recognition)
Guideline for New Language Requests
If you want to request a new language support, a PR with 1 following files are needed:
- In folder ppocr/utils/dict, it is necessary to submit the dict text to this path and name it with
{language}_dict.txt
that contains a list of all characters. Please see the format example from other files in that folder.
If your language has unique elements, please tell me in advance within any way, such as useful links, wikipedia and so on.
More details, please refer to Multilingual OCR Development Plan.
This project is released under Apache 2.0 license
RetroSearch is an open source project built by @garambo
| Open a GitHub Issue
Search and Browse the WWW like it's 1997 | Search results from DuckDuckGo
HTML:
3.2
| Encoding:
UTF-8
| Version:
0.7.4