diff --git a/ja/tutorials/partner-nodes/google/gemini-omni-flash.mdx b/ja/tutorials/partner-nodes/google/gemini-omni-flash.mdx index 91d9e34bd..2d28d27aa 100644 --- a/ja/tutorials/partner-nodes/google/gemini-omni-flash.mdx +++ b/ja/tutorials/partner-nodes/google/gemini-omni-flash.mdx @@ -1,94 +1,179 @@ --- title: "Gemini Omni Flash: 会話型ビデオ生成" -description: "Gemini Omni Flashは、Googleのマルチモーダルビデオモデルです。パートナーノードを通じてComfyUIで利用でき、自然言語でビデオを生成・編集できます。" +description: "Gemini Omni Flash 1.1は、Googleのマルチモーダルビデオモデルです。パートナーノードを通じてComfyUIで利用でき、自然言語でビデオを生成・編集できます。" sidebarTitle: "Gemini Omni Flash" -translationSourceHash: 30b833fb +translationSourceHash: 89b8db47 translationFrom: tutorials/partner-nodes/google/gemini-omni-flash.mdx translationBlockHashes: - "_intro": 5d6a39ac - "What Gemini Omni Flash is good at": 7de9d95d - "Workflows": 753af4cb - "Get started": 64517938 + "_intro": ea742af9 + "What Gemini Omni Flash 1.1 is good at": 8dd798c0 + "Available workflows": 03f73f1a + "Get started": 02809da9 --- import ReqHint from "/snippets/ja/tutorials/partner-nodes/req-hint.mdx"; import UpdateReminder from "/snippets/ja/tutorials/update-reminder.mdx"; -Gemini Omni Flashは、Google DeepMindの高品質でコスト効率の高いビデオ生成および会話型編集モデルです。Google I/O 2026でGemini Omniファミリーの一部として初めて発表され、Geminiのマルチモーダル推論とネイティブビデオ作成を組み合わせ、開発者が自然な会話を通じてビデオを生成、編集、リミックスできるようにします。 - - +Gemini Omni Flashは、Google DeepMindの会話型ビデオ生成・編集モデルです。Google I/O 2026でGemini Omniファミリーの一部として初めて発表されました。Geminiのマルチモーダル推論とネイティブビデオ作成を組み合わせ、自然言語でやりたいことを記述し、参照画像やビデオを添付すると、同期されたオーディオ付きのクリップを生成します。現在のバージョンGemini Omni Flash 1.1は2026年8月27日に一般提供(GA)され、キーフレーム補間、360p/4K出力オプション、シーン拡張が追加されました。 -## Gemini Omni Flash が得意なこと +## Gemini Omni Flash 1.1 が得意なこと - **会話型ビデオ編集**: 自然言語でビデオを編集・調整: キャラクターの交換、シーンの再照明、アングルの変更、オブジェクトの追加・削除を、オリジナルのオーディオとビデオトラックを維持したまま実行できます -- **マルチモーダル入力**: テキスト、画像、ビデオ入力を結合して生成をガイドします。すべてのビデオ出力に同期したオーディオをネイティブに生成します +- **マルチモーダル入力**: テキスト、画像(最大14枚)、ビデオ(最大3本、各10秒)を組み合わせて生成をガイドします。すべての出力にネイティブのオーディオトラックが付きます +- **キーフレーム補間**: `image_to_video` タスクでは、先頭フレームとオプションの末尾フレームを添付すると、その間の映像をモデルが生成します +- **参照画像からビデオへ**: `` のようなタグで参照画像を役割にバインドすると、画像内のキャラクター、製品、オブジェクトが新しいシーンに登場します。画像自体はフレームとして使用されません +- **シーン拡張**: `extend` タスクはクリップに最大10秒の新しい映像を追加します。直近10秒のコンテキストを読み取ってキャラクターと動きの一貫性を保ち、累計で約40秒まで拡張できます +- **解像度コントロール**: 360pで下書きし(コストは720pの約3分の1)、その後720p、1080p、4Kで再レンダリングします。16:9と9:16のアスペクト比に対応しています - **世界知識とシミュレーション**: 物理の理解と、Gemini の歴史・科学・文化的文脈に関する知識を組み合わせ、フォトリアリズムを超えた意味のあるストーリーテリングを可能にします - **テキストとアクションの同期**: 読みやすいテキストとグラフィックスをビデオに直接レンダリングし、動的タイポグラフィを画面上の動きに同期させます -- **料金**: ビデオ出力1秒あたり $0.10。Veo 3.1 Fast の料金に準じます -## ワークフロー + +Gemini Omni Flash 1.1 のワークフローには ComfyUI 0.34.2 以降が必要です。ノードのモデルドロップダウンで **Omni Flash 1.1** を選択するとGAモデルを使用できます。**Omni Flash** オプションはプレビューモデルを実行します。このモデルは2026年9月30日に廃止される予定です。 + + +## 利用可能なワークフロー + +### テキストから動画へ(Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_t2v} -### テキストから動画へ +Gemini Omni Flash 1.1を使用して、自然言語のプロンプトから映画的なビデオを生成します。プロンプトで希望の長さ(3〜10秒)を直接記述し、ノードでアスペクト比と出力解像度を選択します: 16:9または9:16、低コストの下書きには360p、最終レンダリングには最大4K。すべてのクリップに生成されたオーディオトラックが含まれます。 + +Gemini Omni Flash 1.1 テキストからビデオへのワークフロープレビュー + + - + Comfy Cloudで開く - - JSONをダウンロードするか、テンプレートライブラリで「Gemini Omni Flash」を検索 + + JSONをダウンロードするか、テンプレートライブラリで「Gemini Omni 1.1」を検索 + + + +### 画像から動画へ(Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_i2v} + +Gemini Omni Flash 1.1で画像をアニメーション化します。`image_to_video` タスクでは、添付した1枚目の画像が先頭フレームになり、オプションの2枚目の画像が末尾フレームになります: モデルがその間の映像を生成するため、カメラのオービット、ズームトランジション、ループクリップが予測しやすくなります。 + +Gemini Omni Flash 1.1 画像からビデオへのワークフロープレビュー + + + + + + Comfy Cloudで開く + + + JSONをダウンロードするか、テンプレートライブラリで「Gemini Omni 1.1」を検索 + + + +
+入力素材 + +ワークフローを試すには、以下のサンプル入力画像をダウンロードしてください: + + + + サンプル入力画像をダウンロード +
-![Gemini Omni Flash テキストからビデオへのワークフロープレビュー](https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/templates/api_google_gemini_omni_flash_t2v-1.webp) +### 参照画像からビデオへ(Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_r2v} -自然言語のプロンプトから映画的なビデオを生成します。テキストによる説明を、世界認識に基づく動き、照明、音声を備えたビデオ出力に変換します。ソーシャルメディアコンテンツ作成、迅速なビデオプロトタイピング、反復的なビジュアルストーリーテリングに最適です。 +最大14枚の参照画像から、特定の被写体を取り込んだビデオを生成します。参照モードでは、`` のようなタグで各画像を役割にバインドし、プロンプト内でそのタグを参照します: 画像内のキャラクター、製品、オブジェクトがシーンに登場し、画像自体はフレームとして使用されません。キャラクター参照とスタイル参照を組み合わせると、ブランド一貫性のあるコンテンツにできます。 -### 画像から動画へ +Gemini Omni Flash 1.1 参照画像からビデオへのワークフロープレビュー + + - + Comfy Cloudで開く - - JSONをダウンロードするか、テンプレートライブラリで「Gemini Omni Flash」を検索 + + JSONをダウンロードするか、テンプレートライブラリで「Gemini Omni 1.1」を検索 - - このワークフローで使用する例の入力画像を取得 + + +
+入力素材 + +ワークフローを試すには、以下のサンプル入力画像をダウンロードしてください: + + + + 1つ目のサンプル参照画像をダウンロード - - 2つ目のサンプル入力画像を取得 + + 2つ目のサンプル参照画像をダウンロード +
-![Gemini Omni Flash 画像からビデオへのワークフロープレビュー](https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/templates/api_google_gemini_omni_flash_i2v-1.webp) +### ビデオ編集(Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_edit} -Gemini Omni Flashを使用して2枚の画像からビデオを生成します。自然言語のプロンプトを解釈して、再生時間とアスペクト比を制御します。短いブランドクリップ、ダイナミックなソーシャルメディアコンテンツ、会話型プロンプトによる反復的なビデオ編集に最適です。 +Gemini Omni Flash 1.1を使用して、自然言語でビデオを編集します。`edit` タスクでは、ノードは入力ビデオを1本だけ受け取り(10秒以内)、指示に基づいて書き換えます: 背景の交換、シーンのスタイル変更、要素の追加・削除。`edit` と `extend` タスクは入力ビデオのアスペクト比を維持します。シンプルなプロンプトが最も効果的です。「他はそのままに」と添えると一貫性が最大になります。 -### ビデオ編集 +Gemini Omni Flash 1.1 ビデオ編集ワークフロープレビュー + + - + Comfy Cloudで開く - - JSONをダウンロードするか、テンプレートライブラリで「Gemini Omni Flash」を検索 + + JSONをダウンロードするか、テンプレートライブラリで「Gemini Omni 1.1」を検索 - - このワークフローで使用する例の入力ビデオを取得 + + +
+入力素材 + +ワークフローを試すには、このサンプル入力ビデオをダウンロードしてください: + + + + サンプル入力ビデオをダウンロード +
+ +### ビデオ拡張(Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_extend} -![Gemini Omni Flash ビデオ編集ワークフロープレビュー](https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/templates/api_google_gemini_omni_flash_video_edit-1.webp) +`extend` タスクで既存のビデオを1ステップあたり最大10秒拡張し、合計約40秒のストーリーを構築できます。モデルは直近10秒のコンテキストを分析し、シーンが中断した箇所から継続する際にキャラクター、動き、オーディオの一貫性を保ちます。参照画像を添付して、ストーリーの途中で新しいキャラクターを登場させることもできます。拡張はクリップの末尾にのみ新しいコンテンツを追加しますが、トランジションを自然にするためモデルが元の末尾フレームを修正する場合があります。 -Gemini Omni Flashを使用して、自然言語でビデオを編集します。単一の入力ビデオを、説明文に基づいて1つの編集済み出力に変換します。プロンプトで再生時間とアスペクト比を指定します。ソーシャルメディアでの素早いリミックス、映画的なシーン調整、反復的なビデオの洗練に最適です。 +Gemini Omni Flash 1.1 ビデオ拡張ワークフロープレビュー + + + + Comfy Cloudで開く + + + JSONをダウンロードするか、テンプレートライブラリで「Gemini Omni 1.1」を検索 + + + +
+入力素材 + +ワークフローを試すには、このサンプル入力ビデオをダウンロードしてください: + + + + サンプル入力ビデオをダウンロード + + +
## はじめる -1. ComfyUIを最新バージョンにアップデートする +1. ComfyUIを最新バージョンにアップデートする(Omni Flash 1.1 ワークフローには0.34.2以降が必要) 2. キャンバスをダブルクリックし、「Gemini Omni Flash」ノードを検索する 3. またはテンプレートライブラリから既製のワークフローを使用する 4. 入力タイプ(テキスト、画像、ビデオ)に合ったワークフローを選択する diff --git a/ko/tutorials/partner-nodes/google/gemini-omni-flash.mdx b/ko/tutorials/partner-nodes/google/gemini-omni-flash.mdx index e7812e750..4c539c467 100644 --- a/ko/tutorials/partner-nodes/google/gemini-omni-flash.mdx +++ b/ko/tutorials/partner-nodes/google/gemini-omni-flash.mdx @@ -1,95 +1,179 @@ --- title: "Gemini Omni Flash: 대화형 비디오 생성" -description: "Gemini Omni Flash를 사용하여 자연어로 비디오를 생성하고 편집하세요. Google의 멀티모달 비디오 모델로, ComfyUI에서 파트너 노드를 통해 사용 가능합니다." +description: "Gemini Omni Flash 1.1을 사용하여 자연어로 비디오를 생성하고 편집하세요. Google의 멀티모달 비디오 모델로, ComfyUI에서 파트너 노드를 통해 사용 가능합니다." sidebarTitle: "Gemini Omni Flash" -translationSourceHash: 30b833fb +translationSourceHash: 89b8db47 translationFrom: tutorials/partner-nodes/google/gemini-omni-flash.mdx translationBlockHashes: - "_intro": 5d6a39ac - "What Gemini Omni Flash is good at": 7de9d95d - "Workflows": 753af4cb - "Get started": 64517938 + "_intro": ea742af9 + "What Gemini Omni Flash 1.1 is good at": 8dd798c0 + "Available workflows": 03f73f1a + "Get started": 02809da9 --- import ReqHint from "/snippets/ko/tutorials/partner-nodes/req-hint.mdx"; import UpdateReminder from "/snippets/ko/tutorials/update-reminder.mdx"; -Gemini Omni Flash는 Google DeepMind의 고품질, 비용 효율적인 비디오 생성 및 대화형 편집 모델입니다. Google I/O 2026에서 Gemini Omni 제품군의 일부로 처음 소개되었으며, Gemini의 멀티모달 추론과 네이티브 비디오 생성을 결합하여 개발자가 자연어 대화를 통해 비디오를 생성, 편집 및 리믹스할 수 있도록 합니다. +Gemini Omni Flash는 Google DeepMind의 대화형 비디오 생성 및 편집 모델입니다. Google I/O 2026에서 Gemini Omni 제품군의 일부로 처음 소개되었으며, Gemini의 멀티모달 추론과 네이티브 비디오 생성을 결합합니다. 자연어로 원하는 내용을 설명하고 참조 이미지나 비디오를 첨부하면 동기화된 오디오가 포함된 클립을 생성합니다. 현재 버전인 Gemini Omni Flash 1.1은 2026년 8월 27일 정식 출시(GA)되었으며, 키프레임 보간, 360p/4K 출력 옵션, 장면 확장 기능이 추가되었습니다. - - - -## Gemini Omni Flash가 뛰어난 점 +## Gemini Omni Flash 1.1이 뛰어난 점 - **대화형 비디오 편집**: 자연어를 사용하여 비디오를 다듬고 편집합니다: 캐릭터 교체, 장면 조명 변경, 각도 변경, 원본 오디오 및 비디오 트랙을 유지하면서 객체 추가 또는 제거 -- **멀티모달 입력**: 텍스트, 이미지, 비디오 입력을 결합하여 생성을 안내합니다. 모든 비디오 출력과 함께 동기화된 오디오를 자체적으로 생성합니다 +- **멀티모달 입력**: 텍스트, 이미지(최대 14장), 비디오(최대 3개, 각 10초)를 결합하여 생성을 안내합니다. 모든 출력에 네이티브 오디오 트랙이 포함됩니다 +- **키프레임 보간**: `image_to_video` 작업에서 시작 프레임과 선택적인 종료 프레임을 첨부하면 모델이 그 사이의 영상을 생성합니다 +- **참조 이미지 기반 비디오 생성**: `` 같은 태그로 참조 이미지를 역할에 바인딩하면 이미지 속 캐릭터, 제품, 객체가 새 장면에 등장합니다. 이미지 자체는 프레임으로 사용되지 않습니다 +- **장면 확장**: `extend` 작업은 클립에 최대 10초의 새로운 영상을 추가합니다. 마지막 10초의 컨텍스트를 읽어 캐릭터와 움직임의 일관성을 유지하며, 누적 약 40초까지 확장할 수 있습니다 +- **해상도 제어**: 360p로 초안을 작성한 후(비용은 720p의 약 3분의 1) 720p, 1080p 또는 4K로 다시 렌더링합니다. 16:9 및 9:16 화면 비율을 지원합니다 - **세계 지식 및 시뮬레이션**: 물리 이해와 Gemini의 역사, 과학, 문화적 맥락에 대한 지식을 결합하여 포토리얼리즘을 넘어선 의미 있는 스토리텔링을 가능하게 합니다 - **텍스트 및 동작 동기화**: 읽기 쉬운 텍스트와 그래픽을 비디오에 직접 렌더링하여 키네틱 타이포그래피를 화면 내 움직임과 동기화합니다 -- **가격**: 비디오 출력 1초당 $0.10로, Veo 3.1 Fast 가격과 동일합니다 -## 워크플로 + +Gemini Omni Flash 1.1 워크플로에는 ComfyUI 0.34.2 이상이 필요합니다. 노드의 모델 드롭다운에서 **Omni Flash 1.1**을 선택하면 정식 버전 모델을 사용할 수 있습니다. **Omni Flash** 옵션은 프리뷰 모델을 실행하며, 이 모델은 2026년 9월 30일에 지원이 종료될 예정입니다. + + +## 사용 가능한 워크플로 -### 텍스트 기반 비디오 생성 +### 텍스트 기반 비디오 생성 (Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_t2v} + +Gemini Omni Flash 1.1을 사용하여 자연어 프롬프트로 시네마틱 비디오를 생성합니다. 프롬프트에 원하는 길이(3~10초)를 직접 설명하고, 노드에서 화면 비율과 출력 해상도를 선택합니다: 16:9 또는 9:16, 저비용 초안은 360p, 최종 렌더링은 최대 4K. 모든 클립에 생성된 오디오 트랙이 포함됩니다. + +Gemini Omni Flash 1.1 텍스트 기반 비디오 생성 워크플로 미리보기 + + - + Comfy Cloud에서 열기 - - JSON 다운로드 또는 템플릿 라이브러리에서 "Gemini Omni Flash" 검색 + + JSON 다운로드 또는 템플릿 라이브러리에서 "Gemini Omni 1.1" 검색 -![Gemini Omni Flash 텍스트 기반 비디오 생성 워크플로 미리보기](https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/templates/api_google_gemini_omni_flash_t2v-1.webp) +### 이미지 기반 비디오 생성 (Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_i2v} -자연어 프롬프트로 시네마틱 비디오를 생성합니다. 텍스트 설명을 세계 인식 모션, 조명 및 사운드가 포함된 비디오 출력으로 변환합니다. 소셜 미디어 콘텐츠 생성, 빠른 비디오 프로토타이핑 및 반복적인 시각적 스토리텔링에 이상적입니다. +Gemini Omni Flash 1.1로 이미지에 애니메이션을 적용합니다. `image_to_video` 작업에서 첫 번째 첨부 이미지가 시작 프레임이 되고, 선택적인 두 번째 이미지가 종료 프레임이 됩니다: 모델이 그 사이의 영상을 생성하므로 카메라 오빗, 줌 전환, 루프 클립을 예측 가능하게 만들 수 있습니다. -### 이미지 기반 비디오 생성 +Gemini Omni Flash 1.1 이미지 기반 비디오 생성 워크플로 미리보기 + + - + Comfy Cloud에서 열기 - - JSON 다운로드 또는 템플릿 라이브러리에서 "Gemini Omni Flash" 검색 + + JSON 다운로드 또는 템플릿 라이브러리에서 "Gemini Omni 1.1" 검색 + + + +
+입력 자료 + +워크플로를 사용해 보려면 다음 샘플 입력 이미지를 다운로드하세요: + + + + 샘플 입력 이미지 다운로드 - - 이 워크플로의 예제 입력 이미지 가져오기 + +
+ +### 참조 이미지 기반 비디오 생성 (Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_r2v} + +최대 14장의 참조 이미지에서 특정 피사체를 통합한 비디오를 생성합니다. 참조 모드에서는 `` 같은 태그로 각 이미지를 역할에 바인딩하고 프롬프트에서 해당 태그를 참조합니다: 이미지 속 캐릭터, 제품, 객체가 장면에 등장하며 이미지 자체는 프레임으로 사용되지 않습니다. 캐릭터 참조와 스타일 참조를 결합하면 브랜드 일관성 있는 콘텐츠를 만들 수 있습니다. + +Gemini Omni Flash 1.1 참조 이미지 기반 비디오 생성 워크플로 미리보기 + + + + + + Comfy Cloud에서 열기 - - 두 번째 예제 입력 이미지 가져오기 + + JSON 다운로드 또는 템플릿 라이브러리에서 "Gemini Omni 1.1" 검색 -![Gemini Omni Flash 이미지 기반 비디오 생성 워크플로 미리보기](https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/templates/api_google_gemini_omni_flash_i2v-1.webp) +
+입력 자료 -Gemini Omni Flash를 사용하여 두 이미지로 비디오를 생성합니다. 자연어 프롬프트를 해석하여 재생 시간과 화면 비율을 제어합니다. 짧은 브랜드 클립, 다이나믹한 소셜 미디어 콘텐츠 제작 및 대화형 프롬프트를 통한 반복적인 비디오 편집에 적합합니다. +워크플로를 사용해 보려면 다음 샘플 입력 이미지를 다운로드하세요: -### 비디오 편집 + + + 첫 번째 샘플 참조 이미지 다운로드 + + + 두 번째 샘플 참조 이미지 다운로드 + + +
+ +### 비디오 편집 (Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_edit} + +Gemini Omni Flash 1.1을 사용하여 자연어로 비디오를 편집합니다. `edit` 작업에서 노드는 입력 비디오를 정확히 하나만 받아(10초 이하) 지시에 따라 다시 작성합니다: 배경 교체, 장면 스타일 변경, 요소 추가 또는 제거. `edit` 및 `extend` 작업은 입력 비디오의 화면 비율을 유지합니다. 간단한 프롬프트가 가장 효과적입니다. "나머지는 그대로 유지"를 추가하면 일관성이 최대화됩니다. + +Gemini Omni Flash 1.1 비디오 편집 워크플로 미리보기 + + - + Comfy Cloud에서 열기 - - JSON 다운로드 또는 템플릿 라이브러리에서 "Gemini Omni Flash" 검색 + + JSON 다운로드 또는 템플릿 라이브러리에서 "Gemini Omni 1.1" 검색 - - 이 워크플로의 예제 입력 비디오 가져오기 + + +
+입력 자료 + +워크플로를 테스트하려면 다음 샘플 입력 비디오를 다운로드하세요: + + + + 샘플 입력 비디오 다운로드 +
+ +### 비디오 확장 (Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_extend} -![Gemini Omni Flash 비디오 편집 워크플로 미리보기](https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/templates/api_google_gemini_omni_flash_video_edit-1.webp) +`extend` 작업으로 기존 비디오를 단계당 최대 10초씩 확장하여 총 약 40초의 스토리를 구축할 수 있습니다. 모델은 마지막 10초의 컨텍스트를 분석하여 장면이 중단된 지점부터 이어질 때 캐릭터, 움직임, 오디오의 일관성을 유지합니다. 참조 이미지를 첨부하여 스토리 중간에 새로운 캐릭터를 등장시킬 수도 있습니다. 확장은 클립 끝에만 새 콘텐츠를 추가하지만, 전환이 자연스럽도록 모델이 마지막 원본 프레임을 수정할 수 있습니다. -Gemini Omni Flash를 사용하여 자연어로 비디오를 편집합니다. 하나의 입력 비디오를 설명 지침에 따라 하나의 편집된 출력으로 변환합니다. 프롬프트에서 재생 시간과 화면 비율을 지정합니다. 빠른 소셜 미디어 리믹스, 시네마틱 장면 조정 및 반복적인 비디오 다듬기에 이상적입니다. +Gemini Omni Flash 1.1 비디오 확장 워크플로 미리보기 + + + + Comfy Cloud에서 열기 + + + JSON 다운로드 또는 템플릿 라이브러리에서 "Gemini Omni 1.1" 검색 + + + +
+입력 자료 + +워크플로를 테스트하려면 다음 샘플 입력 비디오를 다운로드하세요: + + + + 샘플 입력 비디오 다운로드 + + +
## 시작하기 -1. ComfyUI를 최신 버전으로 업데이트하세요. +1. ComfyUI를 최신 버전으로 업데이트하세요 (Omni Flash 1.1 워크플로에는 0.34.2 이상 필요) 2. 캔버스를 더블 클릭하고 "Gemini Omni Flash" 노드를 검색하세요. 3. 또는 템플릿 라이브러리에서 바로 사용할 수 있는 워크플로를 사용하세요. 4. 입력 유형(텍스트, 이미지 또는 비디오)에 맞는 워크플로를 선택하세요. diff --git a/tutorials/partner-nodes/google/gemini-omni-flash.mdx b/tutorials/partner-nodes/google/gemini-omni-flash.mdx index 86a462c6b..6f82c8ee6 100644 --- a/tutorials/partner-nodes/google/gemini-omni-flash.mdx +++ b/tutorials/partner-nodes/google/gemini-omni-flash.mdx @@ -1,84 +1,172 @@ --- title: "Gemini Omni Flash: Conversational Video Generation" -description: "Generate and edit videos through natural language using Gemini Omni Flash, Google's multimodal video model, available in ComfyUI through Partner Nodes" +description: "Generate and edit videos through natural language using Gemini Omni Flash 1.1, Google's multimodal video model, available in ComfyUI through Partner Nodes" sidebarTitle: "Gemini Omni Flash" --- + import ReqHint from "/snippets/tutorials/partner-nodes/req-hint.mdx"; import UpdateReminder from "/snippets/tutorials/update-reminder.mdx"; -Gemini Omni Flash is Google DeepMind's high-quality, cost-efficient video generation and conversational editing model. First introduced at Google I/O 2026 as part of the Gemini Omni family, it combines Gemini's multimodal reasoning with native video creation, enabling developers to generate, edit, and remix videos through natural conversation. +Gemini Omni Flash is Google DeepMind's conversational video generation and editing model, part of the Gemini Omni family first introduced at Google I/O 2026. It combines Gemini's multimodal reasoning with native video creation: you describe what you want in plain language, attach reference images or videos, and the model generates a clip with synchronized audio. The current version, Gemini Omni Flash 1.1, reached general availability on August 27, 2026 and adds keyframe interpolation, 360p/4K output options, and scene extension. -## What Gemini Omni Flash is good at +## What Gemini Omni Flash 1.1 is good at -- **Conversational video editing**: Refine and edit videos using natural language: swap characters, relight scenes, alter angles, add or remove objects while maintaining original audio and video tracks -- **Multimodal input**: Combine text, images, and video inputs to guide generation. Natively generates synchronized audio with every video output +- **Conversational video editing**: Refine and edit videos using natural language: swap characters, relight scenes, alter angles, add or remove objects while maintaining the original audio and video tracks +- **Multimodal input**: Combine text, images (up to 14), and videos (up to 3, 10 seconds each) to guide generation. Every output carries a native audio track +- **Keyframe interpolation**: With the `image_to_video` task, attach a starting frame and an optional ending frame, and the model generates the footage in between +- **Reference to video**: Bind reference images to roles with tags like `` so characters, products, and objects from your images appear in a new scene without the image itself being used as a frame +- **Scene extension**: The `extend` task appends up to 10 seconds of new footage to a clip, reading the last 10 seconds of context for character and motion consistency, up to about 40 seconds cumulatively +- **Resolution control**: Draft at 360p (about one third of the 720p cost), then re-render at 720p, 1080p, or 4K. 16:9 and 9:16 aspect ratios are supported - **World knowledge and simulation**: Combines physics understanding with Gemini's knowledge of history, science, and cultural context, enabling meaningful storytelling beyond photorealism - **Text and action synchronization**: Render legible text and graphics directly into video, syncing kinetic typography with on-screen movements -- **Pricing**: $0.10 per second of video output, matching Veo 3.1 Fast pricing -## Workflows + +The Gemini Omni Flash 1.1 workflows require ComfyUI 0.34.2 or later. In the node, select **Omni Flash 1.1** from the model dropdown to use the GA model; the **Omni Flash** option runs the preview model, which is scheduled to retire on September 30, 2026. + + +## Available workflows + +### Text to Video (Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_t2v} + +Generate cinematic video from natural language prompts with Gemini Omni Flash 1.1. Describe the desired length (3 to 10 seconds) directly in the prompt, and pick the aspect ratio and output resolution in the node: 16:9 or 9:16, 360p for cheap drafts, up to 4K for final renders. Every clip includes a generated audio track. -### Text to Video +Gemini Omni Flash 1.1 Text to Video workflow preview + + - + Open in Comfy Cloud - - Download JSON or search "Gemini Omni Flash" in Template Library + + Download JSON or search "Gemini Omni 1.1" in Template Library -![Gemini Omni Flash Text to Video workflow preview](https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/templates/api_google_gemini_omni_flash_t2v-1.webp) +### Image to Video (Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_i2v} + +Animate an image with Gemini Omni Flash 1.1. With the `image_to_video` task, the first attached image becomes the starting frame and an optional second image becomes the ending frame: the model generates the footage in between, which makes camera orbits, zoom transitions, and looping clips predictable. -Generate cinematic video from natural language prompts. Transform text descriptions into video output with world-aware motion, lighting, and sound. Ideal for social media content creation, rapid video prototyping, and iterative visual storytelling. +Gemini Omni Flash 1.1 Image to Video workflow preview -### Image to Video + - + Open in Comfy Cloud - - Download JSON or search "Gemini Omni Flash" in Template Library + + Download JSON or search "Gemini Omni 1.1" in Template Library - - Get the example input image for this workflow + + +
+Input materials + +Download this sample input image to try the workflow: + + + + Download sample input image - - Get the second example input image + +
+ +### Reference to Video (Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_r2v} + +Generate video that incorporates specific subjects from up to 14 reference images. In reference mode, bind each image to a role with tags like `` and refer to the tags in your prompt: characters, products, and objects from your images appear in the scene while the image itself is never used as a frame. Combine character references with style references for brand-consistent content. + +Gemini Omni Flash 1.1 Reference to Video workflow preview + + + + + + Open in Comfy Cloud + + + Download JSON or search "Gemini Omni 1.1" in Template Library -![Gemini Omni Flash Image to Video workflow preview](https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/templates/api_google_gemini_omni_flash_i2v-1.webp) +
+Input materials -Generate a video from two images using Gemini Omni Flash. Interpret natural language prompts to control duration and aspect ratio. Perfect for creating short brand clips, dynamic social media content, and iterative video edits through conversational prompting. +Download these sample input images to try the workflow: -### Video Edit + + + Download the first sample reference image + + + Download the second sample reference image + + +
+ +### Video Edit (Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_edit} + +Edit videos with natural language using Gemini Omni Flash 1.1. With the `edit` task, the node takes exactly one input video (10 seconds or less) and rewrites it based on your instructions: swap backgrounds, restyle scenes, add or remove elements. The `edit` and `extend` tasks keep the aspect ratio of the input video. Simple prompts work best; adding "keep everything else the same" maximizes consistency. + +Gemini Omni Flash 1.1 Video Edit workflow preview + + - + Open in Comfy Cloud - - Download JSON or search "Gemini Omni Flash" in Template Library + + Download JSON or search "Gemini Omni 1.1" in Template Library + + + +
+Input materials + +Download this sample input video to try the workflow: + + + + Download sample input video + + +
+ +### Video Extend (Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_extend} + +Extend an existing video by up to 10 seconds per step with the `extend` task, building stories up to about 40 seconds total. The model analyzes the last 10 seconds of context, keeping characters, motion, and audio coherent as the scene continues from where it left off. Optionally attach reference images to introduce new characters mid-story. Extension appends new content only at the end of the clip, but the model may revise the final source frames to make the transition seamless. + +Gemini Omni Flash 1.1 Video Extend workflow preview + + + + Open in Comfy Cloud - - Get the example input video for this workflow + + Download JSON or search "Gemini Omni 1.1" in Template Library -![Gemini Omni Flash Video Edit workflow preview](https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/templates/api_google_gemini_omni_flash_video_edit-1.webp) +
+Input materials + +Download this sample input video to try the workflow: -Edit videos with natural language using Gemini Omni Flash. Transform a single input video into one edited output based on your descriptive instructions. Specify the duration and aspect ratio in your prompt. Ideal for quick social media remixes, cinematic scene adjustments, and iterative video refinements. + + + Download sample input video + + +
## Get started -1. Update ComfyUI to the latest version +1. Update ComfyUI to the latest version (0.34.2 or later for the Omni Flash 1.1 workflows) 2. Double-click the canvas and search for "Gemini Omni Flash" nodes 3. Or go to the Template Library to use the ready-to-go workflows 4. Choose the workflow that matches your input type (text, image, or video) diff --git a/zh/tutorials/partner-nodes/google/gemini-omni-flash.mdx b/zh/tutorials/partner-nodes/google/gemini-omni-flash.mdx index 9cb0ed200..5b76c52a7 100644 --- a/zh/tutorials/partner-nodes/google/gemini-omni-flash.mdx +++ b/zh/tutorials/partner-nodes/google/gemini-omni-flash.mdx @@ -1,92 +1,179 @@ --- title: "Gemini Omni Flash:对话式视频生成" -description: "通过合作节点在 ComfyUI 中使用 Google 的多模态视频模型 Gemini Omni Flash,以自然语言生成和编辑视频" +description: "通过合作节点在 ComfyUI 中使用 Google 的多模态视频模型 Gemini Omni Flash 1.1,以自然语言生成和编辑视频" sidebarTitle: "Gemini Omni Flash" -translationSourceHash: 30b833fb +translationSourceHash: 89b8db47 translationFrom: tutorials/partner-nodes/google/gemini-omni-flash.mdx translationBlockHashes: - "_intro": 5d6a39ac - "What Gemini Omni Flash is good at": 7de9d95d - "Workflows": 753af4cb - "Get started": 64517938 + "_intro": ea742af9 + "What Gemini Omni Flash 1.1 is good at": 8dd798c0 + "Available workflows": 03f73f1a + "Get started": 02809da9 --- import ReqHint from "/snippets/zh/tutorials/partner-nodes/req-hint.mdx"; import UpdateReminder from "/snippets/zh/tutorials/update-reminder.mdx"; -Gemini Omni Flash 是 Google DeepMind 推出的高质量、经济高效的视频生成与对话式编辑模型。该模型于 Google I/O 2026 作为 Gemini Omni 家族成员首次亮相,将 Gemini 的多模态推理能力与原生的视频创建功能结合,使开发者能够通过自然对话生成、编辑和重新混合视频。 +Gemini Omni Flash 是 Google DeepMind 推出的对话式视频生成与编辑模型,于 Google I/O 2026 作为 Gemini Omni 家族成员首次亮相。它将 Gemini 的多模态推理能力与原生视频创建功能结合:你用自然语言描述想要的内容,附上参考图像或视频,模型即可生成带有同步音频的视频片段。当前版本 Gemini Omni Flash 1.1 已于 2026 年 8 月 27 日正式发布,新增了关键帧插值、360p/4K 输出选项和场景扩展功能。 -## Gemini Omni Flash 的优势 +## Gemini Omni Flash 1.1 的优势 - **对话式视频编辑**:使用自然语言精修和编辑视频:替换角色、重新布光、改变角度、添加或移除对象,同时保留原始音轨和视频轨道 -- **多模态输入**:组合文本、图像和视频输入来引导生成。原生为每个视频输出生成同步音频 +- **多模态输入**:组合文本、图像(最多 14 张)和视频(最多 3 个,每个 10 秒)来引导生成。每个输出都带有原生音频轨道 +- **关键帧插值**:使用 `image_to_video` 任务时,附上起始帧和可选的结束帧,模型会生成两者之间的画面 +- **参考图生视频**:使用 `` 等标签将参考图像绑定到角色,让图像中的角色、产品和对象出现在新场景中,而图像本身不会被用作画面帧 +- **场景扩展**:`extend` 任务为视频追加最多 10 秒的新画面,通过读取最后 10 秒的上下文保持角色和运动的一致性,累计可扩展至约 40 秒 +- **分辨率控制**:先以 360p 起草(成本约为 720p 的三分之一),再以 720p、1080p 或 4K 重新渲染。支持 16:9 和 9:16 画面比例 - **世界知识与模拟**:将物理理解与 Gemini 在历史、科学和文化背景方面的知识相结合,实现超越照片写实的有意义叙事 - **文本与动作同步**:将清晰可读的文本和图形直接渲染到视频中,使动态文字与画面中的动作保持同步 -- **价格**:视频输出每秒 $0.10,与 Veo 3.1 Fast 的定价一致 -## 工作流 + +Gemini Omni Flash 1.1 工作流需要 ComfyUI 0.34.2 或更高版本。在节点中,从模型下拉菜单选择 **Omni Flash 1.1** 即可使用正式版模型;**Omni Flash** 选项运行的是预览版模型,该模型计划于 2026 年 9 月 30 日退役。 + + +## 可用工作流 + +### 文本转视频(Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_t2v} + +使用 Gemini Omni Flash 1.1 根据自然语言提示生成电影级视频。在提示中直接描述所需时长(3 到 10 秒),并在节点中选择画面比例和输出分辨率:16:9 或 9:16,360p 用于低成本草稿,最高可到 4K 用于最终渲染。每个片段都包含生成的音频轨道。 + +Gemini Omni Flash 1.1 文本转视频工作流预览 + + + + + + 在 Comfy Cloud 中打开 + + + 下载 JSON,或在模板库中搜索“Gemini Omni 1.1” + + + +### 图像转视频(Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_i2v} -### 文本转视频 +使用 Gemini Omni Flash 1.1 让图像动起来。使用 `image_to_video` 任务时,第一张附带的图像成为起始帧,可选的第二张图像成为结束帧:模型生成两者之间的画面,让镜头环绕、缩放过渡和循环片段变得可预测。 + +Gemini Omni Flash 1.1 图像转视频工作流预览 + + - + 在 Comfy Cloud 中打开 - - 下载 JSON,或在模板库中搜索“Gemini Omni Flash” + + 下载 JSON,或在模板库中搜索“Gemini Omni 1.1” -![Gemini Omni Flash 文本转视频工作流预览](https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/templates/api_google_gemini_omni_flash_t2v-1.webp) +
+输入素材 + +下载以下示例输入图像以试用工作流: + + + + 下载示例输入图像 + + +
+ +### 参考图生视频(Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_r2v} + +使用最多 14 张参考图像生成融入特定主体的视频。在参考模式下,使用 `` 等标签将每张图像绑定到角色,并在提示中引用这些标签:图像中的角色、产品和对象会出现在场景中,而图像本身不会被用作画面帧。将角色参考与风格参考结合使用,可获得品牌一致性内容。 -根据自然语言提示生成电影级视频。将文本描述转换为具有世界感知的运动、光照和声音的视频输出。非常适合社交媒体内容创作、快速视频原型制作以及迭代式视觉叙事。 +Gemini Omni Flash 1.1 参考图生视频工作流预览 -### 图像转视频 + - + 在 Comfy Cloud 中打开 - - 下载 JSON,或在模板库中搜索“Gemini Omni Flash” + + 下载 JSON,或在模板库中搜索“Gemini Omni 1.1” - - 获取此工作流的示例输入图像 + + +
+输入素材 + +下载以下示例输入图像以试用工作流: + + + + 下载第一张示例参考图像 - - 获取第二张示例输入图像 + + 下载第二张示例参考图像 +
+ +### 视频编辑(Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_edit} -![Gemini Omni Flash 图像转视频工作流预览](https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/templates/api_google_gemini_omni_flash_i2v-1.webp) +使用 Gemini Omni Flash 1.1 以自然语言编辑视频。使用 `edit` 任务时,节点接收一个输入视频(10 秒以内),并根据你的指令重写它:更换背景、重塑场景风格、添加或移除元素。`edit` 和 `extend` 任务会保持输入视频的画面比例。简单的提示效果最好;加上“其余部分保持不变”可以最大化一致性。 -使用 Gemini Omni Flash 从两张图像生成视频。解释自然语言提示以控制时长和画面比例。非常适合制作简短品牌剪辑、动态社交媒体内容,以及通过对话式提示进行迭代视频编辑。 +Gemini Omni Flash 1.1 视频编辑工作流预览 -### 视频编辑 + - + 在 Comfy Cloud 中打开 - - 下载 JSON,或在模板库中搜索“Gemini Omni Flash” + + 下载 JSON,或在模板库中搜索“Gemini Omni 1.1” + + + +
+输入素材 + +下载以下示例输入视频以试用工作流: + + + + 下载示例输入视频 + + +
+ +### 视频扩展(Omni Flash 1.1) {#api_google_gemini_omni_flash_1_1_extend} + +使用 `extend` 任务为现有视频每步扩展最多 10 秒,可构建总长约 40 秒的故事。模型会分析最后 10 秒的上下文,在场景从上次中断处继续时保持角色、运动和音频的连贯。可以选择附上参考图像,在故事中途引入新角色。扩展只在片段末尾追加新内容,但模型可能会修正末尾的原始帧以使过渡更流畅。 + +Gemini Omni Flash 1.1 视频扩展工作流预览 + + + + 在 Comfy Cloud 中打开 - - 获取此工作流的示例输入视频 + + 下载 JSON,或在模板库中搜索“Gemini Omni 1.1” -![Gemini Omni Flash 视频编辑工作流预览](https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/templates/api_google_gemini_omni_flash_video_edit-1.webp) +
+输入素材 + +下载以下示例输入视频以试用工作流: -使用 Gemini Omni Flash 以自然语言编辑视频。根据描述性指令将单个输入视频转换为经过编辑的输出。在提示中指定时长和画面比例。非常适合快速社交媒体混剪、电影场景调整以及迭代视频精修。 + + + 下载示例输入视频 + + +
## 开始使用 -1. 将 ComfyUI 更新到最新版本 +1. 将 ComfyUI 更新到最新版本(Omni Flash 1.1 工作流需要 0.34.2 或更高版本) 2. 双击画布,搜索“Gemini Omni Flash”节点 3. 或者进入模板库,使用现成的工作流 4. 选择与输入类型(文本、图像或视频)匹配的工作流