While auditing a local copy of RELEASE v2025.1.9 against the live corpus, I found a set of versions that are catalogued in the public metadata but whose text files cannot be retrieved.
What I did
I compared the release metadata (metadata/OpenITI_metadata_2025-1-9.tsv, 14,107 rows) against the daily-regenerated kitab-metadata-automation/output/OpenITI_Github_clone_metadata_light.csv (13,364 rows, generated 2026-09-06). Forty-seven Arabic versions appear in the live metadata but not in the release. Of those:
- 3 are upstream URI corrections of texts already in the release, e.g.
0321AbuJacfarTahawi.Mukhtasar.… -> 0321Tahawi.Mukhtasar.Kraken210826162339-ara1
- 1 is genuinely new and publicly retrievable (
0375AnonymousTranslator.TarikhCalamUrusiyus.MAB260526-ara1)
- 43 carry a
url field pointing to OpenITI/9001AH
The issue
https://github.com/OpenITI/9001AH returns 404 for an anonymous user, and so does every raw.githubusercontent.com/OpenITI/9001AH/... path in the metadata. For comparison, github.com/OpenITI/8001AH returns 200, so this is not a client-side problem.
Forty-five rows in the live metadata reference 9001AH. Two of them are in fact in the release -- 0400Anonymous.AlladhiFarisWaRajil.NLIAG000190-ara1 and 0400Anonymous.SayadcufAlanManFihi.NLIAG000191-ara1 -- so 9001AH appears to be a staging area from which some texts are published and others held back.
I appreciate that the withholding is intentional: the RELEASE .gitignore carries
#Libraries not to be put into the corpus:
**Noorlib*
and 39 of the 43 are Noorlib-derived (the other four are ER004, WG001, WG20210922 and Kraken210714194926).
Suggestion
Since the rights position is deliberate, the problem is not the withholding but that the metadata advertises URLs that cannot resolve, which makes an automated completeness check report a corpus as incomplete when it is in fact correct. Two options that would help downstream users:
- Mark these rows in the metadata -- a
status or access value, or simply an empty url -- so consumers can distinguish "restricted" from "missing".
- Document in the release notes that 9001AH-hosted versions are catalogued but not distributed, and say whether access can be requested.
If access can be requested for scholarly use, I would be glad to know the procedure. I work on al-Mawardi and on Ismaili and Shi'i theories of the imamate, and a large part of this set -- five works of al-Qadi al-Nu'man, al-Kirmani's Rahat al-'aql, al-Sijistani, al-Shahrastani, al-Tusi's Masari' al-musari', and al-Jahshiyari's Kitab al-Wuzara' wa-l-kuttab -- is directly relevant. I would use them privately, not redistribute them, and cite by version URI.
Happy to attach the full list of 43 URIs if useful.
While auditing a local copy of RELEASE v2025.1.9 against the live corpus, I found a set of versions that are catalogued in the public metadata but whose text files cannot be retrieved.
What I did
I compared the release metadata (
metadata/OpenITI_metadata_2025-1-9.tsv, 14,107 rows) against the daily-regeneratedkitab-metadata-automation/output/OpenITI_Github_clone_metadata_light.csv(13,364 rows, generated 2026-09-06). Forty-seven Arabic versions appear in the live metadata but not in the release. Of those:0321AbuJacfarTahawi.Mukhtasar.…->0321Tahawi.Mukhtasar.Kraken210826162339-ara10375AnonymousTranslator.TarikhCalamUrusiyus.MAB260526-ara1)urlfield pointing toOpenITI/9001AHThe issue
https://github.com/OpenITI/9001AHreturns 404 for an anonymous user, and so does everyraw.githubusercontent.com/OpenITI/9001AH/...path in the metadata. For comparison,github.com/OpenITI/8001AHreturns 200, so this is not a client-side problem.Forty-five rows in the live metadata reference 9001AH. Two of them are in fact in the release --
0400Anonymous.AlladhiFarisWaRajil.NLIAG000190-ara1and0400Anonymous.SayadcufAlanManFihi.NLIAG000191-ara1-- so 9001AH appears to be a staging area from which some texts are published and others held back.I appreciate that the withholding is intentional: the RELEASE
.gitignorecarriesand 39 of the 43 are Noorlib-derived (the other four are
ER004,WG001,WG20210922andKraken210714194926).Suggestion
Since the rights position is deliberate, the problem is not the withholding but that the metadata advertises URLs that cannot resolve, which makes an automated completeness check report a corpus as incomplete when it is in fact correct. Two options that would help downstream users:
statusoraccessvalue, or simply an emptyurl-- so consumers can distinguish "restricted" from "missing".If access can be requested for scholarly use, I would be glad to know the procedure. I work on al-Mawardi and on Ismaili and Shi'i theories of the imamate, and a large part of this set -- five works of al-Qadi al-Nu'man, al-Kirmani's Rahat al-'aql, al-Sijistani, al-Shahrastani, al-Tusi's Masari' al-musari', and al-Jahshiyari's Kitab al-Wuzara' wa-l-kuttab -- is directly relevant. I would use them privately, not redistribute them, and cite by version URI.
Happy to attach the full list of 43 URIs if useful.