NLTK: JVM argument injection bypass via per-call options in the NLTK Stanford wrappers (incomplete fix of CVE-2026-12841)
Critical9.8CVE-2026-79675 · Published Sep 1, 2026 · updated Sep 10, 2026
Affected versions
| Package | Affected | Fixed in |
|---|---|---|
| nltk PyPI | < 3.10.3 | 3.10.3 |
Details and references
## Vulnerability The fix for CVE-2026-12841 (CWE-88, JVM argument injection) added `_validate_java_options()` to block dangerous JVM flags such as `-agentlib`, `-agentpath`, `-javaagent`, `-Xrunjdwp`, and `@argfile` references. However, the validation is only applied when setting global options via `config_java()`. The `java()` function's per-call `options` parameter -- added by PR #3683 (CVE-2026-12615 fix) -- passes options directly to `subprocess.Popen` without calling `_validate_java_options()`. All four Stanford Java wrapper classes accept user-supplied `java_options` and route them through the unvalidated per-call path, bypassing the CVE-2026-12841 fix entirely. ## Root Cause In `nltk/internals.py`, the `java()` function (line 128) accepts an `options` keyword argument. When `options` is not None, it is converted to a list and prepended to the JVM command (lines 211-217) without any validation: ```python # nltk/internals.py, lines 211-217 (HEAD) if options is None: java_options = _java_options # validated by config_java() else: if isinstance(options, str): options = options.split() java_options = list(options) # NO validation cmd = [_java_bin] + java_options + cmd ``` Compare with `config_java()` (line 92) which does validate: ```python # nltk/internals.py, lines 122-123 _validate_java_options(options) _java_options[:] = options ``` The four affected wrapper classes store user-supplied `java_options` without validation and pass them through the unvalidated per-call path: 1. `GenericStanfordParser` (`nltk/parse/stanford.py`): constructor parameter at line 39, stored at line 78, passed at lines 247 and 256 2. `StanfordTagger` (`nltk/tag/stanford.py`): constructor parameter at line 51, stored at line 79, passed at line 118 3. `StanfordTokenizer` (`nltk/tokenize/stanford.py`): constructor parameter at line 43, stored at line 66, passed at line 109 4. `StanfordSegmenter` (`nltk/tokenize/stanford_segmenter.py`): constructor parameter at line 68, stored at line 117, passed at line 337 ## Proof of Concept ```python from nltk.internals import config_java, java, _validate_java_options # 1. The global config_java() path correctly blocks dangerous flags: try: config_java(options=["-agentpath:/tmp/evil.so"]) except ValueError as e: print(f"config_java blocked: {e}") # blocked as expected # 2. The per-call options path does NOT block them: # (Would execute if Java were installed) # java(["SomeClass"], classpath=".", options=["-agentpath:/tmp/evil.so"]) # This passes "-agentpath:/tmp/evil.so" directly to subprocess.Popen # 3. Stanford wrapper classes pass through without validation: # from nltk.parse.stanford import StanfordParser # parser = StanfordParser(java_options="-agentpath:/tmp/evil.so") # parser.parse(...) # dangerous flag reaches JVM # Verify the gap directly: dangerous_opts = ["-agentpath:/tmp/evil.so"] try: _validate_java_options(dangerous_opts) print("Would have been caught") except ValueError: print("Correctly rejected by _validate_java_options()") # But java() itself never calls _validate_java_options(): import inspect source = inspect.getsource(java) assert "_validate_java_options" not in source, "java() does not validate options" print("Confirmed: java() does not call _validate_java_options()") ``` ## Impact An attacker who controls the `java_options` parameter to any NLTK Stanford wrapper class can inject arbitrary JVM flags, including: - `-agentpath:/path/to/malicious.so` -- loads a native agent, achieving arbitrary code execution - `-javaagent:/path/to/malicious.jar` -- loads a Java agent for bytecode manipulation - `-agentlib:jdwp=transport=dt_socket,server=y,address=*:5005` -- enables remote debugging, allowing remote code execution - `@/path/to/argfile` -- expands an argument file, which can smuggle any of the above This is exploitable in scenarios where NLTK is deployed as a service and `java_options` is derived from user input, configura
- CVSS 3.1
- CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H
- Severity from
- GitHub (reviewed advisory)
- Weakness
- CWE-88
- Also known as
- CVE-2026-79675, PYSEC-2026-3749
- github.com/nltk/nltk/security/advisories/GHSA-m4rf-3fr8-xwx3
- nvd.nist.gov/vuln/detail/CVE-2026-79675
- github.com/nltk/nltk/commit/8fa9650b6009aacfdebbc33d2a08d32c0858ea6c
- github.com/nltk/nltk
- github.com/nltk/nltk/releases/tag/v3.10.3
- www.vulncheck.com/advisories/nltk-before-jvm-argument-injection-via-per-call-options
More NLTK advisories
All NLTK| Date | Advisory | Severity | Fixed in |
|---|---|---|---|
| Sep 1 | NLTK: Uncontrolled search path when invoking the Graphviz 'dot' binary CVE-2026-78680High7.8fixed in 3.10.3 | High7.8 | 3.10.3 |
| Sep 2 | NLTK: Uncontrolled recursion in nltk.featstruct.FeatStructReader causes unhandled RecursionError (DoS) via deeply nested feature-structure input CVE-2026-81724Medium5.3fixed in 3.10.3 | Medium5.3 | 3.10.3 |
| Sep 2 | NLTK: Uncontrolled resource consumption in RecursiveDescentParser via ambiguous or left-recursive grammars CVE-2026-12876Mediumfixed in 3.10.3 | Medium | 3.10.3 |
| Sep 2 | NLTK: Quadratic CPU Exhaustion in `XMLCorpusView._read_xml_fragment()` CVE-2026-81723Medium3.7fixed in 3.10.3 | Medium3.7 | 3.10.3 |
| Sep 2 | NLTK: Model-artifact APIs bypass pathsec and touch files outside allowed roots CVE-2026-81726High7.0no fix yet | High7.0 | No fix yet |
| Sep 2 | NLTK: Downloader.download follows hardlinks and overwrites outside-root files CVE-2026-81727Medium7.1fixed in 3.10.3 | Medium7.1 | 3.10.3 |