Skip to content

fix: stream VCF_TO_CSV instead of retaining every record - #331

Merged
kubranarci merged 1 commit into
nf-core:devfrom
adamrtalbot:fix/89-stream-vcf-to-csv-nfcore
Oct 5, 2026
Merged

kubranarci merged 1 commit into
nf-core:devfrom
adamrtalbot:fix/89-stream-vcf-to-csv-nfcore

Conversation

@adamrtalbot

Copy link
Copy Markdown

Description

bin/vcf_to_csv.py kept every (row, info_dict) until EOF, so peak memory grew with the number of records. It now streams the VCF once to learn the header and the five optional-column flags (SUPP_VEC, SUPP, type_inferred, SVTYPE, SVLEN), reopens the same path, and writes one CSV row at a time.

The CLI, parse_info_field, extract_gt_from_sample, and the CSV dialect are unchanged. Output is byte-identical to the previous converter. No workflow, resource, or .nf-core.yml changes.

Related: adamrtalbot/hap-rs#89

PR checklist

  • This comment contains a description of changes (with reason).
  • If you've fixed a bug or added code that should be tested, add tests!
    Local check: byte-identical against the previous converter, and nf-test test modules/local/custom/vcf_to_csv/tests/main.nf.test --profile=+docker (2 passed, existing snapshot unchanged). No new test file; the module snapshot already covers the output contract.
  • If you've added a new tool - have you followed the pipeline conventions in the contribution docs
  • If necessary, also make a PR on the nf-core/variantbenchmarking branch on the nf-core/test-datasets repository.
  • Make sure your code lints (nf-core pipelines lint).
  • Ensure the test suite passes (nextflow run . -profile test,docker --outdir <OUTDIR>).
  • Check for unexpected warnings in debug mode (nextflow run . -profile debug,test,docker --outdir <OUTDIR>).
  • Usage Documentation in docs/usage.md is updated.
  • Output Documentation in docs/output.md is updated.
  • CHANGELOG.md is updated.
  • README.md is updated (including new tool citations and authors/contributors).

Docs, changelog, and README are unchanged because the CSV contract is the same.

Discover optional INFO columns in one pass, then reopen the VCF and write
one CSV row at a time. Peak memory no longer grows with record count, and
the CSV contract stays byte-identical.

Fixes adamrtalbot/hap-rs#89
adamrtalbot added a commit to adamrtalbot/variantbenchmarking that referenced this pull request Oct 4, 2026
bin/vcf_to_csv.py: two streaming passes, bounded memory. Fixes adamrtalbot/hap-rs#89. Keep nf-core#88 open until executor kill reason is verified. Upstream: nf-core#331.
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown

nf-core pipelines lint overall result: Passed ✅ ⚠️

Posted for pipeline commit c6784fe

+| ✅ 305 tests passed       |+
#| ❔   3 tests were ignored |#
#| ❔   1 tests had warnings |#
!| ❗   1 tests had warnings |!
Details

❗ Test warnings:

❔ Tests ignored:

  • files_unchanged - File ignored due to lint config: docs/images/nf-core-variantbenchmarking_logo_light.png
  • files_unchanged - File ignored due to lint config: docs/images/nf-core-variantbenchmarking_logo_dark.png
  • files_unchanged - File ignored due to lint config: .gitignore or .prettierignore

❔ Tests fixed:

✅ Tests passed:

Run details

  • nf-core/tools version 4.1.0
  • Run at 2026-10-04 08:12:51

@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown

❌ nf-test failed with latest Nextflow version

Note

Tests with Nextflow's latest version failed but it will not cause a CI workflow failure.
Please check if the failure is expected with newer (edge-)releases of Nextflow or if it needs fixing.

  • ❌ docker | latest-everything | Shard 6/7

See the full run for details.

@kubranarci kubranarci left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@VictorDidier is the developer of this script. I guess the changes are good to go but I will ask for his approval.

@kubranarci

Copy link
Copy Markdown
Contributor

Thanks @adamrtalbot for the fix!

@pontushojer

Copy link
Copy Markdown

Thanks @adamrtalbot make a PR for this so fast!

I think this two-pass approach is the best solution if it is important to maintain the same exact CSV output (maybe @kubranarci or @VictorDidier could give input on this).

If not, have you considered a single-pass solution where all extra columns (e.g. SVTYPE) are represented in the CSV, regardless of their presence in the VCF? This would be faster and the output CSV would have the same format regardless of input.

@adamrtalbot

Copy link
Copy Markdown
Author

Thanks @adamrtalbot make a PR for this so fast!

I think this two-pass approach is the best solution if it is important to maintain the same exact CSV output (maybe @kubranarci or @VictorDidier could give input on this).

If not, have you considered a single-pass solution where all extra columns (e.g. SVTYPE) are represented in the CSV, regardless of their presence in the VCF? This would be faster and the output CSV would have the same format regardless of input.

Honestly, I just made it match as closely as I could! My gut feeling is to merge now to fix the expanding memory, then open a fresh PR for any further improvements.

@adamrtalbot

Copy link
Copy Markdown
Author

I have to admit I didn't see the open issue, I found the issue myself while analysing long read VCFs.

@kubranarci

Copy link
Copy Markdown
Contributor

Thanks @adamrtalbot make a PR for this so fast!

I think this two-pass approach is the best solution if it is important to maintain the same exact CSV output (maybe @kubranarci or @VictorDidier could give input on this).

If not, have you considered a single-pass solution where all extra columns (e.g. SVTYPE) are represented in the CSV, regardless of their presence in the VCF? This would be faster and the output CSV would have the same format regardless of input.

this module doesnt have a csv output https://github.com/nf-core/variantbenchmarking/blob/dev/modules/local/plots/svlen_dist/main.nf, it uses CSV input. As long as it produces understandable plots I am eager to test.

@kubranarci

Copy link
Copy Markdown
Contributor

I have to admit I didn't see the open issue, I found the issue myself while analysing long read VCFs.

I guess there were many people affected from this. My tests on big-long vcf files is limited to be honest.

@pontushojer

Copy link
Copy Markdown

Honestly, I just made it match as closely as I could! My gut feeling is to merge now to fix the expanding memory, then open a fresh PR for any further improvements.

Sound good to me, just wanted to lift this as something to consider. Happy to go with this solution :)

@kubranarci

Copy link
Copy Markdown
Contributor

ok then lets merge this to dev, please let me know if it works fine for you as well @pontushojer

@kubranarci
kubranarci merged commit ee85a2e into nf-core:dev Oct 5, 2026
22 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

PLOTS_UPSET out of memory Label change for process VCF_TO_CSV

3 participants