[python][자료구조] "pdf 취소선 검출"  How to identify strike-out text from PDF files using Python

I would like to extract only the strike-out text from a .pdf file. I have tried the below code, it is working with a sample pdf file I have. But it is not working with another pdf file which I think is a scanned one. Is there any standard way to extract only strike-out text from a pdf file using python? Any help would be really appreciated.

This is the code I was using:

from pydoc import doc
from pdf2docx import parse
from typing import Tuple
from docx import Document

def convert_pdf2docx(input_file: str, output_file: str, pages: Tuple = None):
    """Converts pdf to docx"""
    if pages:
        pages = [int(i) for i in list(pages) if i.isnumeric()]
    result = parse(pdf_file=input_file,
                   docx_with_path=output_file, pages=pages)
    summary = {
        "File": input_file, "Pages": str(pages), "Output File": output_file
    }

if __name__ == "__main__":
    pdf_file = 'D:/AWS practice/sample_striken_out.pdf'
    doc_file = 'D:/AWS practice/sample_striken_out.docx'
    convert_pdf2docx(pdf_file, doc_file)
    document = Document(doc_file)
    with open('D:/AWS practice/sample_striken_out.txt', 'w') as f:
        for p in document.paragraphs:
            for run in p.runs:
                if not run.font.strike:
                    f.write(run.text)
                    print(run.text)
            f.write('\n')

Note: I am converting PDF to DOCX first and then trying to identify the strike-out text. This code is working with a sample file. But it is not working with the scanned pdf file. The pdf to doc conversion is taking place, but the strike-through detection does not.

 

 

[출처] https://stackoverflow.com/questions/72601927/how-to-identify-strike-out-text-from-pdf-files-using-python

경축! 아무것도 안하여 에스천사게임즈가 새로운 모습으로 재오픈 하였습니다.
어린이용이며, 설치가 필요없는 브라우저 게임입니다.
https://s1004games.com

 

 

 

 

 

본 웹사이트는 광고를 포함하고 있습니다.
광고 클릭에서 발생하는 수익금은 모두 웹사이트 서버의 유지 및 관리, 그리고 기술 콘텐츠 향상을 위해 쓰여집니다.
번호 제목 글쓴이 날짜 조회 수
66 [Django] 장고 CORS 크로스 도메인 이슈 졸리운_곰 2023.12.17 504
65 [django] CORS(Cross Origin Resource Sharing) in Django 졸리운_곰 2023.12.17 421
64 [django] Forbidden (CSRF cookie not set.) 오류 해결하기 file 졸리운_곰 2023.12.17 346
63 [Django] Django 서버를 domain forwarding 시 : Django nginx Refused to display in a frame because it set 'X-Frame-Options' to 'SAMEORIGIN' 졸리운_곰 2023.12.17 404
62 [Django] Django SSLServer 설정 및 구동 방법 file 졸리운_곰 2023.12.16 288
61 [Django][Django restframework] Django REST framework 시작하기 file 졸리운_곰 2023.05.07 365
60 [Django] REST API 로그인 서버 만들기 (2) - DB 연동, 테스트 file 졸리운_곰 2023.05.07 803
59 [Django] REST API 로그인 서버 만들기 (1) - 코드 졸리운_곰 2023.05.07 430
58 [Django] user의 ip address 가지고 오기 졸리운_곰 2023.05.07 461
57 [Django] [Python_Django] You are trying to add a non-nullable field '필드명' to post without a default 해결 졸리운_곰 2023.05.07 325
56 [Django] [Python_Django] 관리자 계정에서 테이블 관리하기 file 졸리운_곰 2023.05.06 397
55 [Django] TIP - 장고 데이터베이스 여러개 사용하기 (Django multidatabase) file 졸리운_곰 2023.05.05 345
54 [Django] 장고 동작 flow에 대해 알아보자! file 졸리운_곰 2023.05.01 455
53 [python django] 내 이해를 도울 Django flow file 졸리운_곰 2023.05.01 502
52 [Django] 개발 Flow & 개발 환경 세팅 file 졸리운_곰 2023.05.01 386
51 [Django] 장고 명령어 기초 정리 졸리운_곰 2023.05.01 439
50 [python django] 장고 자주 쓰는 manage.py 명령 졸리운_곰 2023.05.01 610
49 [python][Django] Python Package Trends: Visualize Package Download Stats in Django file 졸리운_곰 2023.03.18 368
48 [python][Django] Django Workflow and Architecture 장고 개발 워크플로우 및 구조 file 졸리운_곰 2023.03.12 338
47 [python][Django] DRF(장고 rest framework)와 REST API서버 개발 file 졸리운_곰 2023.03.12 341
대표 김성준 주소 : 경기 용인 분당수지 U타워 등록번호 : 142-07-27414
통신판매업 신고 : 제2012-용인수지-0185호 출판업 신고 : 수지구청 제 123호 개인정보보호최고책임자 : 김성준 sjkim70@stechstar.com
대표전화 : 010-4589-2193 [fax] 02-6280-1294 COPYRIGHT(C) stechstar.com ALL RIGHTS RESERVED