Compare commits

..
1 Commits
Author SHA1 Message Date
retoor 711c3b4802 Last version 2024-11-22 20:37:42 +01:00
3 changed files with 6 additions and 26 deletions
-20
View File
@@ -1,20 +0,0 @@
name: pdf2text test
run-name: syntax check
on: [push]
jobs:
Compile:
runs-on: ubuntu-latest
steps:
- name: Check out repository code
uses: actions/checkout@v4
- name: List files in the repository
run: |
ls ${{ gitea.workspace }}
- run: echo "Install dependencies."
- run: apt update
- run: apt install python3 python3-pip python3-venv -y
- run: python3 -m pip install -r requirements.txt
- run: echo "Check if starts correcly. Syntax check."
- run: ./pdf2text .
- run: echo "This job's status is ${{ job.status }}."
+3 -6
View File
@@ -3,10 +3,10 @@
I've converted 8gb of PDF's to text in one afternoon on a decade old x270 using this script. Performant enough imho. Try to get 8Gb in your LLM and getting it to actually use it. That's the challenge.
## Convert all PDF's to text
This is an [script](/pdf2text) for converting a batch of PDF's to text for machine learning.
This is an script for converting a batch of PDF's to text for machine learning.
It only has two dependencies:
- `python3`
- `pdf.miner` (python requirement, specified in [requirements.txt](/requirements.txt) file)
- python3
- pdf.miner (python requirement, specified in requirements.txt file)
## Installation
```bash
@@ -22,6 +22,3 @@ source .venv/bin/activate
./pdf2text [source/destination dir]
```
You read that correctly, the source directory is also the destination directory.
## Todo:
Make decent python package so it's installable on system without having to load environment first. Not sure if worth it, it's not something you daily use.
+3
View File
@@ -25,6 +25,9 @@ if not source_path.exists():
raise Exception(f"{source_path.absolute()} does not exist.")
print("This script will convert all your pdf files to txt files in the same directory.")
if input("Continue? [Y/n]: ").strip().lower() in ["n", "no"]:
print("Operation cancelled.")
exit(0)
for file in pathlib.Path(source_path).glob("*.pdf"):