Repository navigation
Support large input texts with more than 250 words #71
Description
Activity
- addedenhancementNew feature or requestNew feature or request
on Sep 22, 2025 Hi @kyteinsky! I'd like to work on this issue.
My plan is to:
- Split long text into chunks at sentence boundaries rather than cutting by word count.
- Handle languages such as Chinese/Japanese/Thai without adding unwanted spaces when joining translated chunks.
- Reuse the existing
[translate_batch()]flow to keep the change minimal.Lines 90 to 94 in f133406
results = self.translator.translate_batch( [input_tokens], batch_type="tokens", **self.config["inference"], ) - Only apply chunking to longer inputs, leaving the existing behavior unchanged for shorter texts.
Please let me know if you'd prefer a different approach.
hello, the plan looks good.
there's also the case of RTL (right-to-left) languages like Arabic and Persian. Maybe conversion from one RTL lang to another is fine but in RTL to LTR, the chunks may need to be reversed when joined.hello, the plan looks good. there's also the case of RTL (right-to-left) languages like Arabic and Persian. Maybe conversion from one RTL lang to another is fine but in RTL to LTR, the chunks may need to be reversed when joined.
Thanks @kyteinsky for the suggestion! can i proceed with the implementation and also test RTL cases such as Arabic/Persian → English to verify the chunk ordering and handle them appropriately when joining the translated chunks.
yeah sure, go right ahead.
I am following the Context Chat installation guide and currently working on Step 2, setting up the Deploy Daemon in the AppAPI Admin settings.
The Deploy Daemon connection test is successful, and I have also registered the daemon successfully. However, when I click "Start the Deploy Test", I get the following error:
Error installing ExApp
I have verified that the Deploy Daemon connection itself is successful, but the ExApp installation fails during the deploy test.
Could you please help me understand what could be causing this error and what logs or configuration details I should check to troubleshoot it?
Attaching the config that i have done:
BTW, here i'm getting check connection success
Thanks!
Hi @kyteinsky , I've opened a PR to address this. The implementation uses sentence-boundary chunking (~80 words per chunk, configurable) with forward-order joining. All chunking parameters are exposed in config.json so they can be tuned per deployment. Tested across 7 language pairs including RTL (Arabic, Persian) and verified all input sections appear in the output.
Let me know if you have any feedback.
Reacted by Anupam KumarI am following the Context Chat installation guide and currently working on Step 2, setting up the Deploy Daemon in the AppAPI Admin settings.
hey, sorry for the delay. Can you open an issue in the https://github.com/nextcloud/app_api repo?
Please include the docker container details of the harp container:docker inspect <harp-container-name/id>and the logsdocker logs <harp-container-name/id>.but I don't think it is required for a dev setup.
PS: would be nice to have an issue in context_chat's repos if there's an issue there.Reacted by Ravi Shankar Kumar- Thanks for the clarification! I’m working with a dev setup, and everything is working fine on my end. Appreciate your help!…On Sun, 20 Sept, 2026, 1:46 pm Anupam Kumar, ***@***.***> wrote: *kyteinsky* left a comment (nextcloud/translate2#71) <#71 (comment)> I am following the Context Chat installation guide and currently working on Step 2, setting up the Deploy Daemon in the AppAPI Admin settings. hey, sorry for the delay. Can you open an issue in the https://github.com/nextcloud/app_api repo? Please include the docker container details of the harp container: docker inspect <harp-container-name/id> and the logs docker logs <harp-container-name/id>. but I don't think it is required for a dev setup. PS: would be nice to have an issue in context_chat's repos if there's an issue there. — Reply to this email directly, view it on GitHub <#71?email_source=notifications&email_token=BEXKIPRC5TSAWZKTK5OWP6D5P6G4VA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKNZUHA3DCNBZHA22M4TFMFZW63VHMNXW23LFNZ2KKZLWMVXHJLDGN5XXIZLSL5RWY2LDNM#issuecomment-5748614985>, or unsubscribe <https://github.com/notifications/unsubscribe-auth/BEXKIPRKKE37DYAF7ZD6BH35P6G4VAVCNFSNUABFKJSXA33TNF2G64TZHM3TKNRZGQ2DIMBRHNEXG43VMU5TGNBUGEYTOOBYHE42C5QC> . Triage notifications, keep track of coding agent tasks and review pull requests on the go with GitHub Mobile for iOS <https://github.com/notifications/mobile/ios/BEXKIPW4ABTPLGOAAK2QRND5P6G4VA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKNZUHA3DCNBZHA22M4TFMFZW63VHMNXW23LFNZ2KKZLWMVXHJKTGN5XXIZLSL5UW64Y> and Android <https://github.com/notifications/mobile/android/BEXKIPS6XX5FAAL5FRWIO635P6G4VA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKNZUHA3DCNBZHA22M4TFMFZW63VHMNXW23LFNZ2KKZLWMVXHJLTGN5XXIZLSL5QW4ZDSN5UWI>. Download it today! You are receiving this because you commented.Message ID: ***@***.***>
How to use GitHub
Feature request
Which Nextcloud Version are you currently using: v32.0.0
Is your feature request related to a problem? Please describe.
Large input texts get cut off at some point in the corresponing translation/output text. This seems to vary based on the target language chosen and the
max_decoding_lengthparam does not help here much even with high values.Describe the solution you'd like
Chunking of the input text, maybe in around 100 words, to keep the translation input chunks small and digestable by the model.
Note: split and join of the texts will need some special care depending on the language of the input text, for different separators, RTL languages and no-space languages.
Describe alternatives you've considered
Split the input text by hand.
Additional context
translate_batchfunction under the hood that we use here now: https://opennmt.net/CTranslate2/python/ctranslate2.Translator.html#ctranslate2.Translator.translate_iterable