Comments (5)
Thanks for reporting. I saw similar stuff when the vertical line was too close to the character.
One thing I come up with is to use the columns
option. Can you try it? See details in: https://tabula-py.readthedocs.io/en/latest/tabula.html#tabula.io.read_pdf
Without having the PDF, this is what I can suggest.
from tabula-py.
hi @chezou,
Thanks for the recommendation! I added "columns=[10.1, 20.2, 30.3]" to the code and the column is no long cut off 👍 Am I
from tabula-py.
Ugh, pressed the wrong button sorry @chezou. My follow-up question is if my syntax is correct. Using "columns=[10.1, 20.2, 30.3]" worked, but for future PDFs, I want to fully understand how to use the option. I know it's supposed to be: X coordinates of column boundaries. Can you please explain or suggest a better explanation from the documentation? Thanks!
from tabula-py.
The description comes from tabula-java's one. I'm not sure what your point is.
https://github.com/tabulapdf/tabula-java#commandline-usage-examples
Example code can be found in this article by @tdpetrou https://www.dunderdata.com/blog/read-trapped-tables-within-pdfs-as-pandas-dataframes
If you think you want to set the same columns option between different PDFs, that is not possible. You need to set the columns option per table.
from tabula-py.
Hi,
Thanks again for the documentation and the article. This clears up a lot for me. Thanks again! :)
from tabula-py.
Related Issues (20)
- tabula.io.read_pdf argument "pandas_options" is being changed inside the function HOT 1
- tabula.io.read_pdf argument "pandas_options" is being changed inside the function HOT 3
- Extracting non tabular data from pdfs, is it possible? HOT 1
- Extracting non-tabular (1-tabula output) data from pdf, is it possible? HOT 3
- Unable to remove error : Got stderr: Picked up _JAVA_OPTIONS: -Djava.awt.headless=true HOT 1
- Unable to remove note in log : Got stderr: Picked up _JAVA_OPTIONS: -Djava.awt.headless=true HOT 1
- Tabula py Ignores an entire column if it's blank and if it does not contain headerd? HOT 1
- tabula-py CalledProcessError: Command '['java', '-Dfile.encoding=UTF8', '-jar', HOT 3
- dont ignore empty columns in tables spanning multiple pages HOT 1
- Try to install tabula-py HOT 1
- Use JPype instead of subprocess HOT 11
- Add a way to set areas for non-existent pages in template HOT 4
- Exception: RuntimeError: java.lang.UnsatisfiedLinkError: HOT 2
- cant install tabula-py on m1 mac vscode. HOT 1
- Support Python 3.12 HOT 5
- Pls add "orientation" parameter to read_pdf HOT 4
- Security vulnerability in tabula-1.0.5-jar-with-dependencies.jar HOT 4
- [BUG] Encoding still being overridden even after fix to #371. HOT 5
- FutureWarning: errors='ignore' is deprecated and will raise in a future version. HOT 3
- Unable to detect table with longer header information HOT 4
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from tabula-py.