01 Oct 18:19

63f1b53

v3.8: Memory management for persistent services, numpy 2.0 support Latest

Latest

Optional memory management for persistent services

Support a new context manager method Language.memory_zone(), to allow long-running services to avoid growing memory usage from cached entries in the Vocab or StringStore. Once the memory zone block ends, spaCy will evict Vocab and StringStore entries that were added during the block, freeing up memory. Doc objects created inside a memory zone block should not be accessed outside the block.

The current implementation disables population of the tokenizer cache inside the memory zone, resulting in some performance impact. The performance difference will likely be negligible if you're running a full pipeline, but if you're only running the tokenizer, it'll be much slower. If this is a problem, you can mitigate it by warming the cache first, by processing the first few batches of text without creating a memory zone. Support for memory zones in the tokenizer will be added in a future update.

The Language.memory_zone() context manager also checks for a memory_zone() method on pipeline components, so that components can perform similar memory management if necessary. None of the built-in components currently require this.

If you component needs to add non-transient entries to the StringStore or Vocab, you can pass the allow_transient=False flag to the Vocab.add() or StringStore.add() components.

Example usage:

import spacy
import json
from pathlib import Path
from typing import Iterator
from collections import Counter
import typer
from spacy.util import minibatch


def texts(path: Path) -> Iterator[str]:
    with path.open("r", encoding="utf8") as file_:
        for line in file_:
            yield json.loads(line)["text"]

def main(jsonl_path: Path) -> None:
    nlp = spacy.load("en_core_web_sm")
    counts = Counter()
    batches = minibatch(texts(jsonl_path), 1000)
    for i, batch in enumerate(batches):
        print("Batch", i)
        with nlp.memory_zone():
            for doc in nlp.pipe(batch):
                for token in doc:
                    counts[token.text] += 1
    for word, count in counts.most_common(100):
        print(count, word)

if __name__ == "__main__":
    typer.run(main)

Numpy v2 compatibility

Numpy 2.0 isn't binary-compatible with numpy v1, so we need to build against one or the other. This release isolates the dependency change and has no other changes, to make things easier if the dependency change causes problems.

This dependency change was previously attempted in version 3.7.6, but dependencies within the v3.7 family of models resulted in some conflicts, and some packages depending on numpy v1 were incompatible with v3.7.6. I've therefore removed the 3.7.6 release and replaced it with this one, which increments the minor version.

Model packages no longer list spacy as a requirement

I've also made a change to the way models are packaged to make it easier to release more quickly. Previously spaCy models specified a versioned requirement on spacy itself. This meant that there was no way to increment the spaCy version and have it work with the existing models, because the models would specify they were only compatible with spacy>=3.7.0,<3.8.0. We have a compatibility table that allows spacy to see which models are compatible, but the models themselves can't know which future versions of spaCy they work with.

I've therefore added a flag --require-parent/--no-require-parent to the spacy package CLI, which controls where the parent package (e.g. spaCy) should be listed as a requirement of the model. --require-parent is the default for v3.8, but this will change to --no-require-parent by default in v4. I've set --no-require-parent for the v3.8 models, so that further changes can be published that don't impact the models, without retraining the models or forcing users to redownload them.

Assets 23

spacy-3.8.2-cp310-cp310-macosx_10_9_x86_64.whl

6.3 MB 2024-10-01T18:09:56Z
spacy-3.8.2-cp310-cp310-macosx_11_0_arm64.whl

6.01 MB 2024-10-01T18:09:56Z
spacy-3.8.2-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl

27.8 MB 2024-10-01T18:09:56Z
spacy-3.8.2-cp310-cp310-musllinux_1_2_x86_64.whl

28.5 MB 2024-10-01T18:09:56Z
spacy-3.8.2-cp310-cp310-win_amd64.whl

11.6 MB 2024-10-01T18:09:56Z
spacy-3.8.2-cp311-cp311-macosx_10_9_x86_64.whl

6.27 MB 2024-10-01T18:09:56Z
spacy-3.8.2-cp311-cp311-macosx_11_0_arm64.whl

5.97 MB 2024-10-01T18:09:56Z
spacy-3.8.2-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl

29.2 MB 2024-10-01T18:09:56Z
spacy-3.8.2-cp311-cp311-musllinux_1_2_x86_64.whl

29.7 MB 2024-10-01T18:09:56Z
spacy-3.8.2-cp311-cp311-win_amd64.whl

11.6 MB 2024-10-01T18:09:56Z
Source code (zip)

2024-10-01T14:49:49Z
Source code (tar.gz)

2024-10-01T14:49:49Z

09 Sep 14:19

github-actions

prerelease-v3.8.0.dev0

a019315

Optional memory management for persistent services Pre-release

Pre-release

If you component needs to add non-transient entries to the StringStore or Vocab, you can pass the allow_transient=False flag to the Vocab.add() or StringStore.add() components.

Example usage:

import spacy
import json
from pathlib import Path
from typing import Iterator
from collections import Counter
import typer
from spacy.util import minibatch


def texts(path: Path) -> Iterator[str]:
    with path.open("r", encoding="utf8") as file_:
        for line in file_:
            yield json.loads(line)["text"]

def main(jsonl_path: Path) -> None:
    nlp = spacy.load("en_core_web_sm")
    counts = Counter()
    batches = minibatch(texts(jsonl_path), 1000)
    for i, batch in enumerate(batches):
        print("Batch", i)
        with nlp.vocab.memory_zone():
            for doc in nlp.pipe(batch):
                for token in doc:
                    counts[token.text] += 1
    for word, count in counts.most_common(100):
        print(count, word)

if __name__ == "__main__":
    typer.run(main)```

Assets 19

20 Aug 10:09

github-actions

prerelease-v3.7.6a

ca5aec1

v3.7.6a: Test pypi release process Pre-release

Pre-release

prerelease-v3.7.6a

Try to import cibuildwheel settings from previous setup

Assets 19

05 Jun 07:57

svlandeg

v3.7.5

a6d0fc3

v3.7.5: Download sanitization, Typer compatibility, and a bugfix for linking gold entities

✨ New features and improvements

Sanitize direct download for spacy download (#13313).
Convert Cython properties to decorator syntax (#13390).
Bump Weasel pin to allow v0.4.x (#13409).
Improvements to the test suite (#13469, #13470).
Bump Typer pin to allow v0.10.0 and above (#13471).
Allow typing-extensions<5.0.0 for Python < 3.8 (#13516).

🔴 Bug fixes

#13400: Fix use_gold_ents behaviour for EntityLinker.

📖 Documentation and examples

Make the file name for code listings stick to the top (#13379).
Update the documentation of MorphAnalysis (#13433).
Typo fixes in the documentation (#13466).

👥 Contributors

@danieldk, @honnibal, @ines, @JoeSchiff, @nokados, @Paillat-dev, @rmitsch, @schorfma, @strickvl, @svlandeg, @ynx0

Contributors

danieldk, strickvl, and 9 other contributors

Assets 2

15 Feb 19:16

danieldk

v3.7.4

bff8725

v3.7.4: New textcat layers and fo/nn language extensions

✨ New features and improvements

Improve NumPy 2.0 compatibility (#13103).
Added language extensions for Faroese and Norwegian Nynorsk (#13116).
Add new TextCatReduce.v1 layer for text classification (#13181).
Add new TextCatParametricAttention.v1 layer for text classification (#13201).
Use build module for creating model packages by default (#13109).
Add support for code loading to the benchmark speed command (#13247).
Extend lexical attributes for English with more numericals (#13106).
Warn about reloading dependencies after downloading models (#13081).

🔴 Bug fixes

#13259, #13304, #13321: Correctness fixes for multiprocessing support in Language.pipe.
#13187: Typing and documentation fixes for Doc.
#13086: Update Tokenizer.explain for special cases with whitespace.
#13068: Fix displaCy span stacking.
#13149: Add spacy.TextCatBOW.v3 to use the fixed SparseLinear layer.

📖 Documentation and examples

Many improvements and updates to the LLM documentation.
Update trf_data examples and the transformer pipeline design section.

👥 Contributors

@adrianeboyd, @danieldk, @evornov, @honnibal, @ines, @lise-brinck, @ridge-kimani, @rmitsch, @shadeMe, @svlandeg

Contributors

danieldk, shadeMe, and 8 other contributors

Assets 2

16 Oct 16:11

adrianeboyd

v3.7.2

a89eae9

v3.7.2: Fixes for APIs and requirements

✨ New features and improvements

Update __all__ fields (#13063).

🔴 Bug fixes

#13035: Remove Pathy requirement.
#13053: Restore spacy.cli.project API.
#13057: Support Any comparisons for Token and Span.

📖 Documentation and examples

Many updates for spacy-llm including Azure OpenAI, PaLM, and Mistral support.
Various documentation corrections.

👥 Contributors

@adrianeboyd, @honnibal, @ines, @rmitsch, @svlandeg

Contributors

adrianeboyd, rmitsch, and 3 other contributors

Assets 2

05 Oct 06:46

adrianeboyd

v3.7.1

9d03660

v3.7.1: Bug fix for spacy.cli module loading

🔴 Bug fixes

Revert lazy loading of CLI module for spacy.info to fix availability of spacy.cli following import spacy (#13040).

👥 Contributors

@adrianeboyd, @honnibal, @ines, @svlandeg

Contributors

adrianeboyd, honnibal, and 2 other contributors

Assets 2

02 Oct 09:10

adrianeboyd

v3.7.0

160e617

v3.7.0: Trained pipelines using Curated Transformers and support for Python 3.12

This release drops support for Python 3.6 and adds support for Python 3.12.

✨ New features and improvements

Add support for Python 3.12 (#12979).
Use the new library Weasel for spaCy projects functionality (#12769).
- All spacy project commands should run as before, just now they're using Weasel under the hood.
- ⚠️ Remote storage is not yet supported for Python 3.12. Use Python 3.11 or earlier for remote storage.
Extend to Thinc v8.2 (#12897).
Extend transformers extra to spacy-transformers v1.3 (#13025).
Support registered vectors (#12492).
Add --spans-key option for CLI evaluation with spacy benchmark accuracy (#12981).
Load the CLI module lazily for spacy.info (#12962).
Add type stubs for spacy.training.example (#12801).
Warn for unsupported pattern keys in dependency matcher (#12928).
Language.replace_listeners: Pass the replaced listener and the tok2vec pipe to the callback in order to support spacy-curated-transformers (#12785).
Always use tqdm with disable=None to disable output in non-interactive environments (#12979).
Language updates:
- Add left and right pointing angle brackets as punctuation to ancient Greek (#12829).
- Update example sentences for Turkish (#12895).
Package setup updates:
- Update NumPy build constraints for NumPy 1.25+ (#12839). For Python 3.9+, it is no longer necessary to set build constraints while building binary wheels.
- Refactor Cython profiling in order to disable profiling for Python 3.12 in the package setup, since Cython does not currently support profiling for Python 3.12 (#12979).

📦 Trained pipelines updates

The transformer-based trf pipelines have been updated to use our new Curated Transformers library through the Thinc model wrappers and pipeline component from spaCy Curated Transformers.

⚠️ Backwards incompatibilities

Drop support for Python 3.6.
Drop mypy checks for Python 3.7.
Remove ray extra.
spacy project has a few backwards incompatibilities due to the transition to the standalone library Weasel, which is not as tightly coupled to spaCy. Weasel produces warnings when it detects older spaCy-specific settings in your environment or project config.
- Support for the spacy_version configuration key has been dropped.
- Support for the check_requirements configuration key has been dropped due to the deprecation of pkg_resources.
- The SPACY_CONFIG_OVERRIDES environment variable is no longer checked. You can set configuration overrides using WEASEL_CONFIG_OVERRIDES.
- Support for SPACY_PROJECT_USE_GIT_VERSION environment variable has been dropped.
- Error codes are now Weasel-specific and do not follow spaCy error codes.

📖 Documentation and examples

New and updated documentation for large language models and spaCy Curated Transformers.
Various documentation corrections and updates.
New additions to the spaCy Universe:
- Hobbit spaCy: NLP for Middle Earth
- rolegal: a spaCy Package for Noisy Romanian Legal Document Processing

👥 Contributors

@adrianeboyd, @bdura, @connorbrinton, @danieldk, @davidberenstein1957, @denizcodeyaa, @eltociear, @evornov, @honnibal, @ines, @jmyerston, @koaning, @magdaaniol, @pdhall99, @ringohoffman, @rmitsch, @senisioi, @shadeMe, @svlandeg, @vinbo8, @wjbmattingly

Contributors

danieldk, shadeMe, and 19 other contributors

Assets 2

08 Aug 14:45

adrianeboyd

v3.6.1

458bc5f

v3.6.1: Support for Pydantic v2, find-function CLI and more

✨ New features and improvements

Allow Pydantic v2 using transitional v1 support (#12888).
Add find-function CLI for finding locations of registered functions (#12757).
Add extra spacy[cuda12x] for cupy-cuda12x (#12890).
Extend tests for init config and train CLI (#12173).
Switch from distutils to setuptools/sysconfig (#12853).

🔴 Bug fixes

#12817: Escape annotated HTML tags in displaCy span renderer.
#12857: Display model's full base version string in incompatibility warning.
#12882: Update <br> tags in displaCy.

📖 Documentation and examples

Various documentation corrections and updates.
New additions to spaCy Universe:
- OdyCy
- SaysWho

👥 Contributors

@adrianeboyd, @afriedman412, @arplusman, @bdura, @connorbrinton, @honnibal, @ines, @it176131, @pmbaumgartner, @rmitsch, @shadeMe, @svlandeg, @thomashacker, @victorialslocum, @x-tabdeveloping

Contributors

shadeMe, connorbrinton, and 13 other contributors

Assets 2

07 Jul 08:32

adrianeboyd

v3.6.0

6fc153a

v3.6.0: New span finder component and pipelines for Slovenian

✨ New features and improvements

NEW: span_finder pipeline component to identify overlapping, unlabeled spans (#12507).
Language updates:
- Add initial support for Malay (#12602).
- Update Latin defaults to support noun chunks, update lexical/tokenizer defaults and add example sentences (#12538).
Add option to return scores separately keyed by component name with spacy evaluate --per-component, Language.evaluate(per_component=True) and Scorer.score(per_component=True) (#12540).
Support custom token/lexeme attribute for vectors (#12625).
Support spancat_singlelabel in spacy debug data CLI (#12749).
Typing updates for PhraseMatcher and SpanGroup (#12642, #12714).

🔴 Bug fixes

#12569: Require that all SpanGroup spans come from the current doc.

📦 Trained pipelines updates

We have added new pipelines for Slovenian that use the trainable lemmatizer and floret vectors.

Package	UPOS	Parser LAS	NER F
`sl_core_news_sm`	96.9	82.1	62.9
`sl_core_news_md`	97.6	84.3	73.5
`sl_core_news_lg`	97.7	84.3	79.0
`sl_core_news_trf`	99.0	91.7	90.0

🙏 Special thanks to @orglce for help with the new pipelines!

The English pipelines have been updated to improve handling of contractions with various apostrophes and to lemmatize "get" as a passive auxiliary.

The Danish pipeline da_core_news_trf has been updated to use vesteinn/DanskBERT with performance improvements across the board.

⚠️ Backwards incompatibilities

SpanGroup spans are now required to be from the same doc. When initializing a SpanGroup, there is a new check to verify that all added spans refer to the current doc. Without this check, it was possible to run into string store or other errors.