# Open source: ocrJob - an ocrmypdf GUI front-end

**URL:** <https://forum.xojo.com/t/open-source-ocrjob-an-ocrmypdf-gui-front-end/75654>\
**Category:** Code Sharing\
**Created:** [May 7, 2023, 4:03pm UTC](https://forum.xojo.com/t/open-source-ocrjob-an-ocrmypdf-gui-front-end/75654 "2023-05-07T16:03:15Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![Georgios\_Poulopoulos](https://forum.xojo.com/user_avatar/forum.xojo.com/georgios_poulopoulos/32/16905_2.png) [@Georgios\_Poulopoulos](https://forum.xojo.com/u/Georgios_Poulopoulos)\
**Post date:** [May 7, 2023, 4:03pm UTC](https://forum.xojo.com/t/open-source-ocrjob-an-ocrmypdf-gui-front-end/75654/1 "2023-05-07T16:03:15Z")

</div>

Hello all,

I’ve open sourced [ocrJob](https://github.com/gregorplop/ocrJob), a GUI application for creating and executing batch OCR jobs using ocrmypdf/Tesseract.

 ![ocrJobSetup](https://forum.xojo.com/uploads/default/original/2X/9/95923be4e791c83deda75090fee4c813663737cd.jpeg)

It is primarily focused on Windows, but I guess it can work on MacOS and Linux. I just don’t have the time to test it on these two other plaforms. Plus, I don’t have a Mac.

It is in its initial stages of development, so if you experience and bugs or inconsistencies, let me know.  
If you also have any suggestions for improvements, or for integrating it into your workflows, I could implement them if they are generalizable enough.

Cheers,

George

---

<div class="post-metadata">

**Author:** ![Julia\_Truchsess](https://forum.xojo.com/user_avatar/forum.xojo.com/julia_truchsess/32/499_2.png) [@Julia\_Truchsess](https://forum.xojo.com/u/Julia_Truchsess)\
**Post date:** [May 7, 2023, 4:14pm UTC](https://forum.xojo.com/t/open-source-ocrjob-an-ocrmypdf-gui-front-end/75654/2 "2023-05-07T16:14:15Z")

</div>

I tried Tesseract for my OCR application but found it was very sensitive to image quality (I’m working with video caps with a lot of variation in target orientation, lighting, etc.). I switched to AWS Textract and while not free, the performance is a quantum leap above anything I could get from Tesseract. It finds every scrap of text in the image even under poor lighting and at crazy angles.

---

<div class="post-metadata">

**Author:** ![Georgios\_Poulopoulos](https://forum.xojo.com/user_avatar/forum.xojo.com/georgios_poulopoulos/32/16905_2.png) [@Georgios\_Poulopoulos](https://forum.xojo.com/u/Georgios_Poulopoulos)\
**Post date:** [May 7, 2023, 5:14pm UTC](https://forum.xojo.com/t/open-source-ocrjob-an-ocrmypdf-gui-front-end/75654/3 "2023-05-07T17:14:05Z")

</div>

That’s nice to know!  
My own experience is that on-premise Tesseract’s prime use case is this:

- Best effort OCR/searchable PDF creation. Not if the OCR output is mission-critical in some way.
- You can’t send the content to someone else’s computer (ie the Cloud)
- The source images are of some controlled quality: text documents from a scanner. No video feeds and the like.

But it’s true, Tessaract output’s quality is noticeably lower than commercial engines. Sometimes, it’s good enough 🙂

---

<div class="post-metadata">

**Author:** ![Emile\_Schwarz](https://forum.xojo.com/letter_avatar_proxy/v4/letter/e/90ced4/32.png) [@Emile\_Schwarz](https://forum.xojo.com/u/Emile_Schwarz)\
**Post date:** [May 8, 2023, 7:21am UTC](https://forum.xojo.com/t/open-source-ocrjob-an-ocrmypdf-gui-front-end/75654/4 "2023-05-08T07:21:58Z")

</div>

OK, on macOS only, but have-you tried LiveText (Ventura) ?

Julia ?

---

<div class="post-metadata">

**Author:** ![Jean-Yves\_Pochez](https://forum.xojo.com/user_avatar/forum.xojo.com/jean-yves_pochez/32/21769_2.png) [@Jean-Yves\_Pochez](https://forum.xojo.com/u/Jean-Yves_Pochez)\
**Post date:** [May 8, 2023, 8:14am UTC](https://forum.xojo.com/t/open-source-ocrjob-an-ocrmypdf-gui-front-end/75654/5 "2023-05-08T08:14:19Z")

</div>

@Emile_Schwarz , how would you interface livetext with your xojo app ???!?

---

<div class="post-metadata">

**Author:** ![Emile\_Schwarz](https://forum.xojo.com/letter_avatar_proxy/v4/letter/e/90ced4/32.png) [@Emile\_Schwarz](https://forum.xojo.com/u/Emile_Schwarz)\
**Post date:** [May 8, 2023, 10:05am UTC](https://forum.xojo.com/t/open-source-ocrjob-an-ocrmypdf-gui-front-end/75654/6 "2023-05-08T10:05:33Z")

</div>

Declares ?

I have read someone have used Quick Look with Xojo, so… maybe ?

---

<div class="post-metadata">

**Author:** ![Robert\_Schofield](https://forum.xojo.com/letter_avatar_proxy/v4/letter/r/eada6e/32.png) [@Robert\_Schofield](https://forum.xojo.com/u/Robert_Schofield)\
**Post date:** [May 8, 2023, 1:10pm UTC](https://forum.xojo.com/t/open-source-ocrjob-an-ocrmypdf-gui-front-end/75654/7 "2023-05-08T13:10:30Z")

</div>

It is part of Visionkit so likely with declares.

> **[Vision | Apple Developer Documentation](https://developer.apple.com/documentation/vision)**
>
> Apply computer vision algorithms to perform a variety of tasks on input images and video.

---

<div class="post-metadata">

**Author:** ![Julia\_Truchsess](https://forum.xojo.com/user_avatar/forum.xojo.com/julia_truchsess/32/499_2.png) [@Julia\_Truchsess](https://forum.xojo.com/u/Julia_Truchsess)\
**Post date:** [May 8, 2023, 5:09pm UTC](https://forum.xojo.com/t/open-source-ocrjob-an-ocrmypdf-gui-front-end/75654/8 "2023-05-08T17:09:46Z")

</div>

My app has to be cross-platform, so Mac-only is a non-starter for me. AWS Textract was pretty easy to set up in an hour or two.

---

<div class="post-metadata">

**Author:** ![Patrice\_C](https://forum.xojo.com/user_avatar/forum.xojo.com/patrice_c/32/14346_2.png) [@Patrice\_C](https://forum.xojo.com/u/Patrice_C)\
**Post date:** [May 9, 2023, 11:31pm UTC](https://forum.xojo.com/t/open-source-ocrjob-an-ocrmypdf-gui-front-end/75654/9 "2023-05-09T23:31:58Z")

</div>

Always easy to do mac development by using a remote mac from [macweb.com](http://macweb.com)

---

<div class="post-metadata">

**Author:** ![Georgios\_Poulopoulos](https://forum.xojo.com/user_avatar/forum.xojo.com/georgios_poulopoulos/32/16905_2.png) [@Georgios\_Poulopoulos](https://forum.xojo.com/u/Georgios_Poulopoulos)\
**Post date:** [May 10, 2023, 3:58pm UTC](https://forum.xojo.com/t/open-source-ocrjob-an-ocrmypdf-gui-front-end/75654/10 "2023-05-10T15:58:37Z")

</div>

ah, you tried it on a Mac and works ok?  
that’s cool! 🙂

---

<div class="post-metadata">

**Author:** ![Georgios\_Poulopoulos](https://forum.xojo.com/user_avatar/forum.xojo.com/georgios_poulopoulos/32/16905_2.png) [@Georgios\_Poulopoulos](https://forum.xojo.com/u/Georgios_Poulopoulos)\
**Post date:** [May 10, 2023, 4:00pm UTC](https://forum.xojo.com/t/open-source-ocrjob-an-ocrmypdf-gui-front-end/75654/11 "2023-05-10T16:00:45Z")

</div>

thanks, will keep it mind! I don’t see myself making any money from developing for the Mac any time soon though 🙂

---

<div class="post-metadata">

**Author:** ![Georgios\_Poulopoulos](https://forum.xojo.com/user_avatar/forum.xojo.com/georgios_poulopoulos/32/16905_2.png) [@Georgios\_Poulopoulos](https://forum.xojo.com/u/Georgios_Poulopoulos)\
**Post date:** [May 14, 2023, 11:22am UTC](https://forum.xojo.com/t/open-source-ocrjob-an-ocrmypdf-gui-front-end/75654/12 "2023-05-14T11:22:41Z")

</div>

**I have a serious bug warning about versions 1.x.x:** the last document in the queue is not getting OCR’ed!  
It’s being fixed in version 2.0.0, along with a big refactoring of the job engine.

Sorry about that, totally slipped through 🫤
