CatchTheTornado

text-extract-api

CatchTheTornado

Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown

AI 简介

这是一个本地化部署的文档内容提取与结构化转换API,支持PDF、Word、PPTX等格式及图像文件的高精度OCR识别与语义解析。核心功能包括:基于EasyOCR和Ollama多模型(如Llama 3.2 Vision、MiniCPM-V)的混合OCR策略;LLM后处理优化文本准确性;自动去除PII敏感信息;输出结构化JSON或语义保留的Markdown;支持异步任务队列(Celery)与Redis缓存。适用于金融合规审查、医疗报告结构化、发票信息抽取、政务文档脱敏等需数据不出内网、强隐私保护的场景。

Python
MIT License
3.2k
Stars
277
Forks
14
Watchers
46
Issues

Star 增长

今日0
近 7 天0
近 30 天0
综合评分59.33
默认分支main