Skip to content

Support large input texts with more than 250 words #71

Description

@kyteinsky

How to use GitHub

  • Please use the 👍 reaction to show that you are interested into the same feature.
  • Please don't comment if you have no relevant information to add. It's just extra noise for everyone subscribed to this issue.
  • Subscribe to receive notifications on status change and new comments.

Feature request

Which Nextcloud Version are you currently using: v32.0.0

Is your feature request related to a problem? Please describe.
Large input texts get cut off at some point in the corresponing translation/output text. This seems to vary based on the target language chosen and the max_decoding_length param does not help here much even with high values.

Describe the solution you'd like
Chunking of the input text, maybe in around 100 words, to keep the translation input chunks small and digestable by the model.
Note: split and join of the texts will need some special care depending on the language of the input text, for different separators, RTL languages and no-space languages.

Describe alternatives you've considered
Split the input text by hand.

Additional context

Activity

  1. RSKKSOFFICIAL commented on Aug 30, 2026

    @RSKKSOFFICIAL

    Hi @kyteinsky! I'd like to work on this issue.

    My plan is to:

    • Split long text into chunks at sentence boundaries rather than cutting by word count.
    • Handle languages such as Chinese/Japanese/Thai without adding unwanted spaces when joining translated chunks.
    • Reuse the existing [translate_batch()] flow to keep the change minimal.
      results = self.translator.translate_batch(
      [input_tokens],
      batch_type="tokens",
      **self.config["inference"],
      )
    • Only apply chunking to longer inputs, leaving the existing behavior unchanged for shorter texts.

    Please let me know if you'd prefer a different approach.

  2. kyteinsky commented on Aug 31, 2026

    @kyteinsky
    ContributorAuthor

    hello, the plan looks good.
    there's also the case of RTL (right-to-left) languages like Arabic and Persian. Maybe conversion from one RTL lang to another is fine but in RTL to LTR, the chunks may need to be reversed when joined.

  3. RSKKSOFFICIAL commented on Sep 1, 2026

    @RSKKSOFFICIAL

    hello, the plan looks good. there's also the case of RTL (right-to-left) languages like Arabic and Persian. Maybe conversion from one RTL lang to another is fine but in RTL to LTR, the chunks may need to be reversed when joined.

    Thanks @kyteinsky for the suggestion! can i proceed with the implementation and also test RTL cases such as Arabic/Persian → English to verify the chunk ordering and handle them appropriately when joining the translated chunks.

  4. kyteinsky commented on Sep 1, 2026

    @kyteinsky
    ContributorAuthor

    yeah sure, go right ahead.

  5. RSKKSOFFICIAL commented on Sep 1, 2026

    @RSKKSOFFICIAL

    I am following the Context Chat installation guide and currently working on Step 2, setting up the Deploy Daemon in the AppAPI Admin settings.

    The Deploy Daemon connection test is successful, and I have also registered the daemon successfully. However, when I click "Start the Deploy Test", I get the following error:

    Error installing ExApp

    Image

    I have verified that the Deploy Daemon connection itself is successful, but the ExApp installation fails during the deploy test.

    Could you please help me understand what could be causing this error and what logs or configuration details I should check to troubleshoot it?

    Attaching the config that i have done:

    Image Image

    BTW, here i'm getting check connection success

    Thanks!

  6. RSKKSOFFICIAL commented on Sep 5, 2026

    @RSKKSOFFICIAL

    Hi @kyteinsky , I've opened a PR to address this. The implementation uses sentence-boundary chunking (~80 words per chunk, configurable) with forward-order joining. All chunking parameters are exposed in config.json so they can be tuned per deployment. Tested across 7 language pairs including RTL (Arabic, Persian) and verified all input sections appear in the output.

    Let me know if you have any feedback.

  7. kyteinsky commented on Sep 20, 2026

    @kyteinsky
    ContributorAuthor

    I am following the Context Chat installation guide and currently working on Step 2, setting up the Deploy Daemon in the AppAPI Admin settings.

    hey, sorry for the delay. Can you open an issue in the https://github.com/nextcloud/app_api repo?
    Please include the docker container details of the harp container: docker inspect <harp-container-name/id> and the logs docker logs <harp-container-name/id>.

    but I don't think it is required for a dev setup.
    PS: would be nice to have an issue in context_chat's repos if there's an issue there.

  8. RSKKSOFFICIAL commented on Sep 20, 2026

    @RSKKSOFFICIAL
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions