Skip to content
EchoMiner

DBT-BUILDER · Group 3 · JSS AHER

Echo reports go in as PDFs.Data comes out as a table.

EchoMiner converts semi-structured echocardiography PDF reports into analysis-ready structured data. Developed under the DBT-BUILDER project at JSS Academy of Higher Education and Research, Mysore.

PDFs per batch
5

PDFs per batch

fields per report
47

fields per report

files kept after download
0

files kept after download

Doppler study · spectral envelope

EF61 %
FS34 %
LVIDd48 mm
LA38 mm
EDV108 ml
MVE0.8 A0.6

Illustrative values. Not from a patient record.

About

What EchoMiner does

Echocardiography reports are written for clinicians to read, not for software to analyse. EchoMiner reads them the way a research assistant would, and returns a table.

A departmental echo archive holds thousands of PDF reports. Each one contains the same measurements in roughly the same places, but as prose and layout rather than as data. Extracting them by hand is the step that stops most retrospective studies before they start.

EchoMiner applies a validated rule-based pipeline to those reports and returns one row per study, with the measurement fields, the narrative findings and the impression lines in separate columns — plus a quality summary of what it read from each file.

Rationale

Why structured echocardiography reports matter

Free text cannot be counted

A cohort described in paragraphs cannot be summarised, stratified or modelled without first being turned into variables.

Manual entry does not scale

Transcribing an archive by hand is slow, and every pass introduces variation that is invisible in the final dataset.

Reproducibility needs provenance

A deterministic pipeline with a recorded version lets a reviewer regenerate the same dataset from the same reports.

Features

What you get

Batch extraction

Upload up to 5 echocardiography PDFs at once and receive one consolidated workbook.

Rule-based and reproducible

Deterministic regular-expression extraction. The same input always produces the same output, and every workbook records the engine version that produced it.

Quality reporting

Every export states pages read, reports detected in the source, rows returned, and any parse warnings, per file.

Analysis-ready output

Measurements arrive as numeric cells with a data dictionary sheet, ready for statistical software.

Nothing retained

Uploaded PDFs and generated files are deleted as soon as your download completes.

Citable

Archived on Zenodo with a DOI. Citation formats are included in every workbook.

Workflow

Five steps, one page

Registration and the tool live on this page. Nothing redirects you elsewhere.

  1. 01

    Register once

    Verify your email with a one-time code. Subsequent visits from the same device skip it.

  2. 02

    Upload reports

    Drop up to 5 PDF reports. Files are checked before anything is read.

  3. 03

    Extract

    The validated pipeline reads each report and maps it to the structured field set.

  4. 04

    Review

    Preview the extracted table and the per-file quality summary in the browser.

  5. 05

    Download and clear

    Take the branded workbook. Your files are removed from the server immediately.

Access

Launch EchoMiner

Register once, verify your email, and the tool opens right here on this page.

Already registered?

The information collected through this portal is used solely for providing access to the EchoMiner research platform and maintaining institutional usage records. User information will be securely stored in accordance with applicable Government of India data protection guidelines and institutional policies. The information will not be shared with any third party except where required by law or institutional policy.

This form is protected by reCAPTCHA; the Google Privacy Policy and Terms of Service apply.

Funding

JSSAHER DBT BUILDER Project

JSS Academy of Higher Education & Research (JSSAHER) has been selected by the Department of Biotechnology (DBT) to implement the prestigious DBT BUILDER (Boost to University Interdisciplinary Life Science Departments for Education and Research) program.

Backed by a ₹5 crore grant over five years, this initiative promotes interdepartmental collaboration to nurture postgraduate talent and build a globally competitive bio-economy.

The project targets the prevention and management of cardiovascular diseases across three emerging research domains:

  1. 1Novel Biomarker and TherapeuticsMetabolic disorders and cardiopulmonary disease.
  2. 2NanotheranosticsAdvancing CVD disease management.
  3. 3Spatial Health Informatics and ManagementThe domain EchoMiner is built in.

Spatial Health Informatics & Management (Group 3)

Group 3 leads the spatial health informatics domain and is the driving force behind the development of AI_EchoMiner — a scalable, Python-based data extraction framework that utilizes regular expressions and pandas to process complex healthcare data efficiently.

JSS Academy of Higher Education and ResearchDepartment of Biotechnology, Government of India

People

Project leadership and development team

Project leadership

  • Dr. Prashant M Vishwanath

    Dean (Research), JSSAHER

    prashantv@jssuni.edu.in
  • Dr. Rajesh Kumar Thimmulappa

    Principal Investigator | Professor, Dept. of Biochemistry, JSS Medical College

    kumar_rt@yahoo.com
  • Dr. Madhu B

    Co-Principal Investigator | Professor & Head, Dept. of Community Medicine, JSS Medical College

    madhub@jssuni.edu.in

Group 3 research & development team

  • Dr. Manjunatha M C

    Assistant Professor, Dept. of Community Medicine, JSS Medical College, Mysuru

    mcmanju1@gmail.com
  • Suraj B M

    Senior Research Fellow, Dept. of Community Medicine, JSS Medical College, Mysuru

    surajbm@jssuni.edu.in

Impact

Research impact

EchoMiner exists to make retrospective echocardiography research feasible at archive scale within the DBT-BUILDER cardiovascular programme.

Usage and output metrics are published here once verified by the project team. Figures are drawn from platform records rather than estimates, and this section stays empty until they are confirmed.

Output

Publications

Publications arising from EchoMiner will be listed here as they appear. The software itself is archived and citable:

B Manjunath, S. (2026). EchoMiner: source code for rule-based NLP extraction from echocardiography PDF reports [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.21281483

https://doi.org/10.5281/zenodo.21281483

Updates

News and updates

No updates have been posted yet. Project announcements, releases and new report-layout support will appear here.

Attribution

How to cite EchoMiner

Acknowledge EchoMiner in any publication, thesis, conference paper, report or scientific communication that uses data generated through this tool.

B Manjunath, S. (2026). EchoMiner: source code for rule-based NLP extraction from echocardiography PDF reports [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.21281483

Questions

Frequently asked questions

What file formats can I upload?

PDF only, up to 5 files per submission and 50 MB in total. Up to 10,000 pages per submission — a quarterly archive usually fits in one. The PDF must contain a text layer — scanned images without OCR cannot be read.

Do I need to verify my email every time?

No. A one-time code is required at first registration. After that, returning from the same browser restores your access silently. A new device or browser asks for one code again.

What happens to the reports I upload?

They are held only while your job runs and are deleted the moment your download completes. Jobs that are abandoned are swept automatically. Usage records such as file counts and timestamps are retained; report content is not.

Is the extraction accurate?

The pipeline is rule-based and deterministic, and it is validated against the report layouts it was built for. Every workbook includes a Quality sheet showing what was read from each file so you can verify the output rather than assume it.

How should I cite EchoMiner?

Use the citation in the How to Cite section, also included in every exported workbook in APA, Vancouver, IEEE, BibTeX and RIS.

Who can use EchoMiner?

Registration is open to researchers. Users are responsible for holding the appropriate ethical approvals for any data they process through the tool.

Contact

Get in touch

Dr. Madhu B

Professor & Head

Co-Principal Investigator

Department of Community Medicine

JSS Medical College

JSS Academy of Higher Education and Research

Mysore, India

madhub@jssuni.edu.in

Data privacy

The information collected through this portal is used solely for providing access to the EchoMiner research platform and maintaining institutional usage records. User information will be securely stored in accordance with applicable Government of India data protection guidelines and institutional policies. The information will not be shared with any third party except where required by law or institutional policy.