Skip to content

VCF 9.0.1 Hangs during VCF Automation Deployment #91

Description

@pcgeek2009

Describe the bug

I am trying to deploy a full stack with the command:
New-HoloDeckInstance -Version '9.0.1.0' -InstanceID 'gray' -WorkloadDomainType 'SharedSSO' -NsxEdgeClusterMgmtDomain -NsxEdgeClusterWkldDomain -DeployVcfAutomation -DeploySupervisor

This works fine until it goes to deploy VCF Automation and ends up showing the error:
10-02-2026 05:57:47 SddcMgmtDomain[2175]: [ERROR] Management Domain deployment failed. Check the logs below for more details
10-02-2026 05:57:47 SddcMgmtDomain[2175]: [ERROR] @{name=Retrieve the status of VCF Automation Deployment request; description=Retrieve the status of VCF Automation Deployment request; status=COMPLETED_WITH_FAILURE; creationTimestamp=02/09/2026 19:54:19; updateTimestamp=02/10/2026 05:56:32; errors=System.Object[]}

The deployment retries and then hangs with the holo router showing:
10-02-2026 12:43:56 SddcMgmtDomain[2175]: [INFO] Current task in progress: Retrieve the status of VCF Automation Deployment request
10-02-2026 12:43:56 SddcMgmtDomain[2175]: [INFO] Management Domain is not ready yet. Sleeping for 5 mins

And has been showing this for over 12 hours.

Reproduction steps

1.New-HoloDeckInstance -Version '9.0.1.0' -InstanceID 'gray' -WorkloadDomainType 'SharedSSO' -NsxEdgeClusterMgmtDomain -NsxEdgeClusterWkldDomain -DeployVcfAutomation -DeploySupervisor
2.
3.
...

Expected behavior

Complete the deployment without issue.

Additional context

This is my second attempt, after deleting all previous VM's, including the Holotrouter and redeploying. Seems to hang as the same point. I had been having some DNS issues, etc. with my previous attempt and thought it maybe related. I deleted everything related to the VCF 9 deployment, including the Holo Router and started over. This time it all worked fine until I hit this step again.

Activity

  1. pcgeek2009 commented on Feb 10, 2026

    @pcgeek2009
    Author
    Image

    I do see this in the deployment and it references the following KB:
    https://knowledge.broadcom.com/external/article/406528/lcmvmsp10002-error-when-deploying-vcf-au.html

    However, since this is a fully automated deployment, I am not sure non-lowercase names plays a part.

  2. dhruv-tyagi-broadcom commented on Feb 10, 2026

    @dhruv-tyagi-broadcom
    Collaborator

    There are a few ways to troubleshoot this.

    1. You can log in to VCF Operations > Fleet Management > LifeCycle. On the right pane > VCF Management > Tasks and check the VCF Automation task for any specific errors.

    2. You can SSH into VCF Installer and look at /var/log/vmware/vcf/domainmanager/domainmanager.log and see if there are any errors there

  3. dhruv-tyagi-broadcom commented on Feb 10, 2026

    @dhruv-tyagi-broadcom
    Collaborator
    Image I do see this in the deployment and it references the following KB: https://knowledge.broadcom.com/external/article/406528/lcmvmsp10002-error-when-deploying-vcf-au.html

    However, since this is a fully automated deployment, I am not sure non-lowercase names plays a part.

    We do not use capital FQDN, so this is not applicable here

  4. pcgeek2009 commented on Feb 10, 2026

    @pcgeek2009
    Author

    .GenericStartVCFTask","finishState":"","errorState":"","properties":{},"uiProperties":{"displayText":"overall fips status collection for vcf","displayKey":"vmf::sm::overallfipsstatuscollectionforvcf"},"nodes":[{"symbolicName":"com.vmware.vrealize.lcm.vcf.plugin.tasks.GenericStartVCFTask","symbolicNameTxt":"overallfipsstatuscollectionforvcf-com.vmware.vrealize.lcm.vcf.plugin.tasks.GenericStartVCFTask","type":"SIMPLE","task":"com.vmware.vrealize.lcm.vcf.plugin.tasks.GenericStartVCFTask","properties":{},"uiProper

    Here is an error I found occurring. I am attaching extended lines of the log

    vcfinstaller error log.txt

    in the logs.

  5. wsalonsonexta commented on Feb 11, 2026

    @wsalonsonexta

    Im having this issue with de VCF deployment step... I used the same cli command described. It stucks.

    11-02-2026 18:14:06 SddcMgmtDomain[8909]: [ERROR] Management Domain deployment failed. Check the logs below for more details
    11-02-2026 18:14:06 SddcMgmtDomain[8909]: [ERROR] @{name=Upload VCF Automation binary to VCF Operations fleet management; description=Upload VCF Automation binary to VCF Operations fleet management; status=COMPLETED_WITH_FAILURE; creationTimestamp=02/11/2026 13:18:34; updateTimestamp=02/11/2026 18:12:18; errors=System.Object[]}

    Image
  6. dhruv-tyagi-broadcom commented on Feb 12, 2026

    @dhruv-tyagi-broadcom
    Collaborator

    .GenericStartVCFTask","finishState":"","errorState":"","properties":{},"uiProperties":{"displayText":"overall fips status collection for vcf","displayKey":"vmf::sm::overallfipsstatuscollectionforvcf"},"nodes":[{"symbolicName":"com.vmware.vrealize.lcm.vcf.plugin.tasks.GenericStartVCFTask","symbolicNameTxt":"overallfipsstatuscollectionforvcf-com.vmware.vrealize.lcm.vcf.plugin.tasks.GenericStartVCFTask","type":"SIMPLE","task":"com.vmware.vrealize.lcm.vcf.plugin.tasks.GenericStartVCFTask","properties":{},"uiProper

    Here is an error I found occurring. I am attaching extended lines of the log

    vcfinstaller error log.txt

    in the logs.

    I don't see anything specific here. Can you try #91 (comment)

  7. dhruv-tyagi-broadcom commented on Feb 12, 2026

    @dhruv-tyagi-broadcom
    Collaborator

    Im having this issue with de VCF deployment step... I used the same cli command described. It stucks.

    11-02-2026 18:14:06 SddcMgmtDomain[8909]: [ERROR] Management Domain deployment failed. Check the logs below for more details 11-02-2026 18:14:06 SddcMgmtDomain[8909]: [ERROR] @{name=Upload VCF Automation binary to VCF Operations fleet management; description=Upload VCF Automation binary to VCF Operations fleet management; status=COMPLETED_WITH_FAILURE; creationTimestamp=02/11/2026 13:18:34; updateTimestamp=02/11/2026 18:12:18; errors=System.Object[]}

    Image

    Can you try step 2 and see what the exact error stack is? #91 (comment)

  8. wsalonsonexta commented on Feb 12, 2026

    @wsalonsonexta
  9. dhruv-tyagi-broadcom commented on Feb 12, 2026

    @dhruv-tyagi-broadcom
    Collaborator

    Here is

    domainmanager.log

    This file does not have the deployment logs. Are there any other domainmanager log files in that folder?

  10. wsalonsonexta commented on Feb 12, 2026

    @wsalonsonexta

    These are the folders and files located in the path. What could be helpfully?

    thank you

    Image
  11. dhruv-tyagi-broadcom commented on Feb 12, 2026

    @dhruv-tyagi-broadcom
    Collaborator

    You will want to untar the domainmanager.2026... files and look at those log files to see what the error is.

    Another option is to retry the deployment from the VCF Installer UI and then the domainmanager.log file will get new logs added including the error log but deployment/failure may take time.

  12. wsalonsonexta commented on Feb 12, 2026

    @wsalonsonexta

    You will want to untar the domainmanager.2026... files and look at those log files to see what the error is.

    Another option is to retry the deployment from the VCF Installer UI and then the domainmanager.log file will get new logs added including the error log but deployment/failure may take time.

    the uncompressed logs don't show any error. Is there any log form other appliance helpfully?

  13. dhruv-tyagi-broadcom commented on Feb 12, 2026

    @dhruv-tyagi-broadcom
    Collaborator

    There are a few ways to troubleshoot this.

    1. You can log in to VCF Operations > Fleet Management > LifeCycle. On the right pane > VCF Management > Tasks and check the VCF Automation task for any specific errors.
    2. You can SSH into VCF Installer and look at /var/log/vmware/vcf/domainmanager/domainmanager.log and see if there are any errors there

    yes, you can follow step 1

  14. 5 remaining items

  15. wsalonsonexta commented on Feb 12, 2026

    @wsalonsonexta

    As you can see in the file shared:

    UPLOAD_BINARY_TO_VCF_OPERATIONS_MANAGEMENT_FAILED Upload binary content /nfs/vmware/vcf/nfs-mount/bundle/fd82e8d9-c3bd-5a1c-9b03-2ae76de6d299/fd82e8d9-c3bd-5a1c-9b03-2ae76de6d299/vmsp-vcfa-combined-9.0.1.0.24965341.tar to VCF Operations fleet management failed
    

    Binary upload for VCFA to VCF Ops is failing. This is not really a Holodeck thing, but a VCF thing.

    Are you using an offline depot or online? If offline depot, can you verify the checksum for the VCFA binary that you've placed in the offline depot and compare it with the checksum from the Broadcom Support Portal value? If they don't match, you will need to fix the binary.

    If that is not the issue, can you share the full domain manager log and not just the grepped Error log to see the stack trace has any more details

    Is the online depot...

    no network or internet restrictions.

    actually y see some timeskew errors

    domainmanager.log

  16. pcgeek2009 commented on Feb 12, 2026

    @pcgeek2009
    Author

    I am using the online repository as well.

  17. dhruv-tyagi-broadcom commented on Feb 13, 2026

    @dhruv-tyagi-broadcom
    Collaborator

    There are a few errors I can see from the logs:

    2026-02-12T16:46:11.956+0000 ERROR [vcf_dm,698de6c2574e89955b42a873ded00ef9,2318] [c.v.e.s.c.v.v.VcfOpsMgmtServiceImpl,dm-exec-9]  Error while uploading content from /nfs/vmware/vcf/nfs-mount/bundle/fd82e8d9-c3bd-5a1c-9b03-2ae76de6d299/fd82e8d9-c3bd-5a1c-9b03-2ae76de6d299/vmsp-vcfa-combined-9.0.1.0.24965341.tar to VCF Operations Management with vmid 6581ab37-c144-478c-868f-c12d3b93458d, received status code 204 NO_CONTENT.
    

    The first error says NO_CONTENT. You can log into VCF Ops CLI and check if the file actually exists in the path or not. This error is visible only once and moves on to the next error which is repeated, so I'm thinking the file was available after thee first error.

    2026-02-12T16:53:27.249+0000 ERROR [vcf_dm,698e03d4b70e726cd551c07e7dd01f17,77e8] [c.v.e.s.c.v.v.VcfOpsMgmtServiceImpl,dm-exec-4]  Error in uploading content from /nfs/vmware/vcf/nfs-mount/bundle/fd82e8d9-c3bd-5a1c-9b03-2ae76de6d299/fd82e8d9-c3bd-5a1c-9b03-2ae76de6d299/vmsp-vcfa-combined-9.0.1.0.24965341.tar to VCF Operations Management with vmid 84dbbdba-bfe4-45b0-857e-26d6652011c0
    org.springframework.web.client.ResourceAccessException: I/O error on POST request for "https://opslcm-a.site-a.cnmx/lcm/crepo/api/content/upload/84dbbdba-bfe4-45b0-857e-26d6652011c0": Connection reset by peer
    
    	Suppressed: java.io.IOException: Cannot write application data on closed/failed TLS connection
    

    In this case the connection was reset by peer. This is seen multiple times and I'm not really sure why.

  18. pcgeek2009 commented on Feb 13, 2026

    @pcgeek2009
    Author

    I am going to restart my deployment and have it re-try. When it finally timed out, the automation appliance deployed and the errors are around the pods. I cannot seem to log into the appliance though. Maybe I manually set the domain wrong.

    Image Image
  19. pcgeek2009 commented on Feb 13, 2026

    @pcgeek2009
    Author

    When I restarted the deployment, the SDDC manager in the lab produced these errors in the GUI when I was watching it. Not sure if this is helpful or not.

    Image
  20. pcgeek2009 commented on Feb 13, 2026

    @pcgeek2009
    Author

    Seems like it has now gotten past that portion of the deployment and continued.

    Image Image
  21. dhruv-tyagi-broadcom commented on Feb 13, 2026

    @dhruv-tyagi-broadcom
    Collaborator

    Great!

  22. pcgeek2009 commented on Feb 15, 2026

    @pcgeek2009
    Author

    So, I have now been able to duplicate this issue. While I had gotten past the error reported above, I started having trouble and failures on the NSX that I reported as bug #94. So, with having previous problems, I ran Remove-HoloDeckInstance [-ResetHoloRouter] and decided to start over. I ran the same deployment command shown above, and all seemed to be going well until I hit the automation section, where it seemed to hang with the same error originally reported. The system is configured with the workaround published in bug #89 and deploying with command "New-HoloDeckInstance -Version '9.0.1.0' -InstanceID 'gray' -WorkloadDomainType 'SharedSSO' -NsxEdgeClusterMgmtDomain -NsxEdgeClusterWkldDomain -DeployVcfAutomation -DeploySupervisor"

    One thing I noticed this time was that when I tried to log into the GUI for the installer, I received a "not authorized" message for the admin@local account. I ended up switching to root and running "sudo systemctl restart commonsvcs" which fixed the login issue. As soon as the deployment times out, I am going to see if running the command again completes the installation of the VCF Automation as before.

  23. pcgeek2009 commented on Feb 15, 2026

    @pcgeek2009
    Author

    So, it kept hangs at the Automation deployment as before. I rebooted both the Holorouter and the app installer appliances. I restarted the deployment, and as shown above, the Automation deployment completed. However, I when I got to the NSX deployment, I started hitting the same error I reported in #94.

  24. pcgeek2009 commented on Feb 22, 2026

    @pcgeek2009
    Author

    I have now deleted the previous lab with and rebuilt with all 9.0.2 packages. This still hangs while deploying the automation, but with somewhat different results. It now hangs on the "Deploy package task". It ran for 15 hours and failed. Below is the error it showed.

    Image
  25. dhruv-tyagi-broadcom commented on Feb 23, 2026

    @dhruv-tyagi-broadcom
    Collaborator

    VCFA deployment generally fails when there is high CPU. But here's how you can check further:

    1. SSH into VCFA using:
    ssh vmware-system-user@auto-a.site-a.vcf.lab
    sudo su
    export KUBECONFIG=/etc/kubernetes/admin.conf
    kubectl get pods -A
    kubectl describe pod <pod-name-that-failed>
    

    and you can look for more error details at the k8s level.

  26. pcgeek2009 commented on Feb 24, 2026

    @pcgeek2009
    Author

    So, restarting from the failed state seemed to have gotten past that point and moved onto the NSX Edge deployment. So, the automation portion is now deployed.

  27. dhruv-tyagi-broadcom commented on Feb 24, 2026

    @dhruv-tyagi-broadcom
    Collaborator

    Great. Moving this to a discussion as we did not find any specific bug in Holodeck, but this is a good guide for troubleshooting VCFA deployment

  28. locked and limited conversation to collaborators on Feb 24, 2026
  29. converted this issue into a discussion #103 on Feb 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions