Job Description
Role & responsibilities
Build AI-powered workflows to extract structured information from insurance policy documents.
Develop pipelines for OCR, PDF parsing, and document preprocessing.
Use LLMs and prompt engineering to extract policy features, coverage limits, waiting periods, exclusions, and demographic information.
Normalize extracted information into predefined JSON/QMS schemas.
Improve extraction accuracy across multiple insurer formats.
Validate extracted information using business rules and confidence scoring.
Write clean, modular, and well-documented Python code.
Collaborate with Product and Engineering teams to improve extraction accuracy.
Good to Have
LangChain / LlamaIndex
OCR tools (Tesseract, PaddleOCR, Azure OCR, Google Vision)
PyMuPDF / pdfplumber
Vector Databases
RAG
Docker
FastAPI / Flask
PostgreSQL / MongoDB
AWS / Azure / GCP