Compare commits

..
6 Commits
Author SHA1 Message Date
retoor 47ca978650 chore: remove interactive confirmation prompt from pdf2text conversion script
pdf2text test / Compile (push) Failing after 13s
2024-11-22 19:55:08 +00:00
retoor 969ef8b270 feat: replace broken shell command with echo in test workflow
The CI test workflow previously attempted to run an invalid shell command
"Check if starts correcly. Syntax check." which would fail. This commit
replaces it with a safe echo statement that prints the same message,
allowing the pipeline to proceed to subsequent steps.
2024-11-22 19:51:10 +00:00
retoor cab4498410 fix: add python3-pip and python3-venv packages to CI workflow install step 2024-11-22 19:48:37 +00:00
retoor ef9cfc2511 feat: add initial CI workflow for automated build and syntax check on push
Create .gitea/workflows/test.yaml with a Compile job that runs on every push,
checking out code, installing python3 dependencies from requirements.txt,
and executing pdf2text to validate syntax and startup behavior on ubuntu-latest.
2024-11-22 19:45:58 +00:00
retoor 27e1bc2b7a fix: convert plain-text references to hyperlinks for pdf2text script and requirements.txt in README 2024-11-22 19:41:31 +00:00
retoor b4227c1adc chore: add pdf2text batch converter script with venv setup and readme 2024-11-22 19:37:42 +00:00
3 changed files with 26 additions and 6 deletions
+20
View File
@@ -0,0 +1,20 @@
name: pdf2text test
run-name: syntax check
on: [push]
jobs:
Compile:
runs-on: ubuntu-latest
steps:
- name: Check out repository code
uses: actions/checkout@v4
- name: List files in the repository
run: |
ls ${{ gitea.workspace }}
- run: echo "Install dependencies."
- run: apt update
- run: apt install python3 python3-pip python3-venv -y
- run: python3 -m pip install -r requirements.txt
- run: echo "Check if starts correcly. Syntax check."
- run: ./pdf2text .
- run: echo "This job's status is ${{ job.status }}."
+6 -3
View File
@@ -3,10 +3,10 @@
I've converted 8gb of PDF's to text in one afternoon on a decade old x270 using this script. Performant enough imho. Try to get 8Gb in your LLM and getting it to actually use it. That's the challenge.
## Convert all PDF's to text
This is an script for converting a batch of PDF's to text for machine learning.
This is an [script](/pdf2text) for converting a batch of PDF's to text for machine learning.
It only has two dependencies:
- python3
- pdf.miner (python requirement, specified in requirements.txt file)
- `python3`
- `pdf.miner` (python requirement, specified in [requirements.txt](/requirements.txt) file)
## Installation
```bash
@@ -22,3 +22,6 @@ source .venv/bin/activate
./pdf2text [source/destination dir]
```
You read that correctly, the source directory is also the destination directory.
## Todo:
Make decent python package so it's installable on system without having to load environment first. Not sure if worth it, it's not something you daily use.
-3
View File
@@ -25,9 +25,6 @@ if not source_path.exists():
raise Exception(f"{source_path.absolute()} does not exist.")
print("This script will convert all your pdf files to txt files in the same directory.")
if input("Continue? [Y/n]: ").strip().lower() in ["n", "no"]:
print("Operation cancelled.")
exit(0)
for file in pathlib.Path(source_path).glob("*.pdf"):