Repository navigation
feat(language_model): add LiteLLM provider for 100+ backings - #1386
Conversation
|
cc @sscargal |
edwinyyyu
left a comment
There was a problem hiding this comment.
Is there a reason not to add litellm as an optional dependency in server pyproject.toml?
| def get_litellm_language_model_conf(self, name: str) -> "LiteLLMLanguageModelConf": | ||
| """Get LiteLLM language model configuration by name.""" | ||
| return self.litellm_language_model_confs[name] | ||
|
|
There was a problem hiding this comment.
The parse and to_yaml_dict needs update too
| max_attempts: int, | ||
| generate_response_call_uuid: object, | ||
| ) -> ChatCompletion | AsyncIterator[object] | object: | ||
| import litellm |
There was a problem hiding this comment.
update the project.toml file to install the module
There was a problem hiding this comment.
Pull request overview
Adds a new server-side LiteLLMLanguageModel provider intended to expand MemMachine’s supported LLM backends by delegating requests to the LiteLLM SDK while reusing the existing OpenAI chat-completions parsing/streaming/tool-call logic.
Changes:
- Introduces
LiteLLMLanguageModel(subclassingOpenAIChatCompletionsLanguageModel) that routes requests vialitellm.acompletion. - Extends language model configuration and the
LanguageModelManagerto register/build/remove LiteLLM-backed models. - Adds unit tests for the LiteLLM adapter behavior (dispatch, kwarg forwarding, retries, and parent parsing).
Reviewed changes
Copilot reviewed 4 out of 4 changed files in this pull request and generated 3 comments.
| File | Description |
|---|---|
| packages/server/src/memmachine_server/common/resource_manager/language_model_manager.py | Wires LiteLLM configs into manager lookups and adds a builder for LiteLLMLanguageModel. |
| packages/server/src/memmachine_server/common/language_model/litellm_language_model.py | New LiteLLM adapter that swaps the request implementation to litellm.acompletion and adds retry logic. |
| packages/server/src/memmachine_server/common/configuration/language_model_conf.py | Adds LiteLLMLanguageModelConf and a litellm_language_model_confs collection plus accessors. |
| packages/server/server_tests/memmachine_server/common/language_model/test_litellm_language_model.py | New unit tests validating LiteLLM dispatch/forwarding/retry behavior and inherited parsing. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| from typing import Any | ||
| from unittest.mock import AsyncMock, MagicMock, patch | ||
|
|
| ret: LanguageModel | None = None | ||
| if name in self.conf.openai_responses_language_model_confs: | ||
| ret = self._build_openai_responses_language_model(name) | ||
| if name in self.conf.openai_chat_completions_language_model_confs: | ||
| ret = self._build_openai_chat_completions_language_model(name) | ||
| if name in self.conf.amazon_bedrock_language_model_confs: | ||
| ret = self._build_amazon_bedrock_language_model(name) | ||
| if name in self.conf.litellm_language_model_confs: | ||
| ret = self._build_litellm_language_model(name) |
| litellm = [ | ||
| "litellm>=1.63.0", | ||
| ] |
| try: | ||
| import litellm | ||
| except ImportError as e: | ||
| raise ImportError( | ||
| "litellm is required for LiteLLMLanguageModel. " | ||
| "Install it with: pip install memmachine-server[litellm]" | ||
| ) from e |
|
Please ensure CI passes. |
sscargal
left a comment
There was a problem hiding this comment.
@RheagalFire, in addition to @edwinyyyu and CoPilot feedback, please sign your commits. Thanks.
711c275 to
6b5e6ee
Compare
Signed-off-by: Aarish Alam <[email protected]>
c34dccb to
d03abf1
Compare
|
Signed the commits. All previous feedback (parse/to_yaml_dict, optional dep in server pyproject.toml, lint/type check fixes) was already addressed in earlier commits -- now consolidated into one clean commit. |
|
Many thanks @RheagalFire . There's a |
|
@sscargal can you approve the CI. |
sscargal
left a comment
There was a problem hiding this comment.
LGTM. Thanks for the new feature.
The lock had drifted far behind: 30 of the 52 direct dependencies were behind their latest release, some by a major version. Refreshing it moves 104 locked packages, adds 6 and removes 2, with no downgrades. Nine are major bumps: complexipy 6 -> 8, cryptography 49 -> 50, fastmcp 3 -> 4 (and fastmcp-slim), mcp 1 -> 2, sentence-transformers 5 -> 6, setuptools 83 -> 84, websockets 16 -> 17, xxhash 3 -> 4. Direct dependencies that moved include fastapi 0.139.0 -> 0.141.1, instructor 1.15.4 -> 1.17.0, langchain-aws 1.6.2 -> 1.7.5, neo4j 6.2.0 -> 6.3.0, openai 2.45.0 -> 2.54.0, qdrant-client 1.18.0 -> 1.19.0, sqlalchemy 2.0.51 -> 2.0.52, ty 0.0.59 -> 0.0.79 and uvicorn 0.51.0 -> 0.52.4. Two specifiers had to change for the refresh to be honest: - litellm was capped at <1.85 when the provider was added (MemMachine#1386, which records no reason for the bound), and dependabot later widened it to <1.86. That range ends inside a stretch of releases -- 1.83.8 through 1.92.x -- that declare Requires-Python >=3.10,<3.14. uv ignores that upper bound and installed 1.85.7 anyway, but pip does not: on Python 3.14, `pip install memmachine-server[litellm]` resolved to 1.83.7, not the 1.85.7 in the lock. 1.93.0 is the first release that supports 3.14, so that is the floor now, and the lock moves to 1.100.0. litellm also requires openai<3, which is what holds the OpenAI SDK on the 2.x line; without the floor the resolver satisfies that by walking litellm backwards to 1.83.0 instead. - boto3-stubs was pinned exactly at 1.43.13 while boto3 resolved to 1.43.90, so the stubs no longer described the installed SDK. Two source changes the new versions require, and one cleanup: - ty 0.0.79 rejects assigning over a bound method, so the metrics-factory wiring test installs its fake engine with monkeypatch.setattr. - `list(vector.flat)` is a list of numpy scalars, not `list[float]`, which the newer numpy typing now reports at the assignment. The replay path converts explicitly. - The dev group listed complexipy twice. Left behind deliberately: nebula5-python (see the cap and its reason), openai 3.x (litellm requires openai<3, and instructor's jiter<0.15 pin holds openai below 3.11 regardless; 3.0 also swaps the transport to httpx2), and ruff, still pinned at 0.15.14 -- 0.16 formats Markdown code blocks, which would reformat six docs, and adds LOG004, which flags four `logger.exception` calls, two of them genuine misuses and two in a helper that is only ever called from an except block. All six ty jobs pass, as do `ruff check`, `ruff format --check`, `uv lock --check` and the unit suite. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01BcPSoGsQjnVJS8A5NkoGZN
The lock had drifted far behind: 30 of the 52 direct dependencies were behind their latest release, some by a major version. Refreshing it moves 104 locked packages, adds 6 and removes 2, with no downgrades. Nine are major bumps: complexipy 6 -> 8, cryptography 49 -> 50, fastmcp 3 -> 4 (and fastmcp-slim), mcp 1 -> 2, sentence-transformers 5 -> 6, setuptools 83 -> 84, websockets 16 -> 17, xxhash 3 -> 4. Direct dependencies that moved include fastapi 0.139.0 -> 0.141.1, instructor 1.15.4 -> 1.17.0, langchain-aws 1.6.2 -> 1.7.5, neo4j 6.2.0 -> 6.3.0, openai 2.45.0 -> 2.54.0, qdrant-client 1.18.0 -> 1.19.0, sqlalchemy 2.0.51 -> 2.0.52, ty 0.0.59 -> 0.0.79 and uvicorn 0.51.0 -> 0.52.4. Two specifiers had to change for the refresh to be honest: - litellm was capped at <1.85 when the provider was added (MemMachine#1386, which records no reason for the bound), and dependabot later widened it to <1.86. That range ends inside a stretch of releases -- 1.83.8 through 1.92.x -- that declare Requires-Python >=3.10,<3.14. uv ignores that upper bound and installed 1.85.7 anyway, but pip does not: on Python 3.14, `pip install memmachine-server[litellm]` resolved to 1.83.7, not the 1.85.7 in the lock. 1.93.0 is the first release that supports 3.14, so that is the floor now, and the lock moves to 1.100.0. litellm also requires openai<3, which is what holds the OpenAI SDK on the 2.x line; without the floor the resolver satisfies that by walking litellm backwards to 1.83.0 instead. - boto3-stubs was pinned exactly at 1.43.13 while boto3 resolved to 1.43.90, so the stubs no longer described the installed SDK. Two source changes the new versions require, and one cleanup: - ty 0.0.79 rejects assigning over a bound method, so the metrics-factory wiring test installs its fake engine with monkeypatch.setattr. - `list(vector.flat)` is a list of numpy scalars, not `list[float]`, which the newer numpy typing now reports at the assignment. The replay path converts explicitly. - The dev group listed complexipy twice. Left behind deliberately: nebula5-python (see the cap and its reason), openai 3.x (litellm requires openai<3, and instructor's jiter<0.15 pin holds openai below 3.11 regardless; 3.0 also swaps the transport to httpx2), and ruff, still pinned at 0.15.14 -- 0.16 formats Markdown code blocks, which would reformat six docs, and adds LOG004, which flags four `logger.exception` calls, two of them genuine misuses and two in a helper that is only ever called from an except block. All six ty jobs pass, as do `ruff check`, `ruff format --check`, `uv lock --check` and the unit suite. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01BcPSoGsQjnVJS8A5NkoGZN
The lock had drifted far behind: 30 of the 52 direct dependencies were behind their latest release, some by a major version. Refreshing it moves 104 locked packages, adds 6 and removes 2, with no downgrades. Nine are major bumps: complexipy 6 -> 8, cryptography 49 -> 50, fastmcp 3 -> 4 (and fastmcp-slim), mcp 1 -> 2, sentence-transformers 5 -> 6, setuptools 83 -> 84, websockets 16 -> 17, xxhash 3 -> 4. Direct dependencies that moved include fastapi 0.139.0 -> 0.141.1, instructor 1.15.4 -> 1.17.0, langchain-aws 1.6.2 -> 1.7.5, neo4j 6.2.0 -> 6.3.0, openai 2.45.0 -> 2.54.0, qdrant-client 1.18.0 -> 1.19.0, sqlalchemy 2.0.51 -> 2.0.52, ty 0.0.59 -> 0.0.79 and uvicorn 0.51.0 -> 0.52.4. Two specifiers had to change for the refresh to be honest: - litellm was capped at <1.85 when the provider was added (MemMachine#1386, which records no reason for the bound), and dependabot later widened it to <1.86. That range ends inside a stretch of releases -- 1.83.8 through 1.92.x -- that declare Requires-Python >=3.10,<3.14. uv ignores that upper bound and installed 1.85.7 anyway, but pip does not: on Python 3.14, `pip install memmachine-server[litellm]` resolved to 1.83.7, not the 1.85.7 in the lock. 1.93.0 is the first release that supports 3.14, so that is the floor now, and the lock moves to 1.100.0. litellm also requires openai<3, which is what holds the OpenAI SDK on the 2.x line; without the floor the resolver satisfies that by walking litellm backwards to 1.83.0 instead. - boto3-stubs was pinned exactly at 1.43.13 while boto3 resolved to 1.43.90, so the stubs no longer described the installed SDK. Two source changes the new versions require, and one cleanup: - ty 0.0.79 rejects assigning over a bound method, so the metrics-factory wiring test installs its fake engine with monkeypatch.setattr. - `list(vector.flat)` is a list of numpy scalars, not `list[float]`, which the newer numpy typing now reports at the assignment. The replay path converts explicitly. - The dev group listed complexipy twice. Left behind deliberately: nebula5-python (see the cap and its reason), openai 3.x (litellm requires openai<3, and instructor's jiter<0.15 pin holds openai below 3.11 regardless; 3.0 also swaps the transport to httpx2), and ruff, still pinned at 0.15.14 -- 0.16 formats Markdown code blocks, which would reformat six docs, and adds LOG004, which flags four `logger.exception` calls, two of them genuine misuses and two in a helper that is only ever called from an except block. All six ty jobs pass, as do `ruff check`, `ruff format --check`, `uv lock --check` and the unit suite. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01BcPSoGsQjnVJS8A5NkoGZN
…edkick) (#1596) * Bring the dependency lock up to date The lock had drifted far behind: 30 of the 52 direct dependencies were behind their latest release, some by a major version. Refreshing it moves 104 locked packages, adds 6 and removes 2, with no downgrades. Nine are major bumps: complexipy 6 -> 8, cryptography 49 -> 50, fastmcp 3 -> 4 (and fastmcp-slim), mcp 1 -> 2, sentence-transformers 5 -> 6, setuptools 83 -> 84, websockets 16 -> 17, xxhash 3 -> 4. Direct dependencies that moved include fastapi 0.139.0 -> 0.141.1, instructor 1.15.4 -> 1.17.0, langchain-aws 1.6.2 -> 1.7.5, neo4j 6.2.0 -> 6.3.0, openai 2.45.0 -> 2.54.0, qdrant-client 1.18.0 -> 1.19.0, sqlalchemy 2.0.51 -> 2.0.52, ty 0.0.59 -> 0.0.79 and uvicorn 0.51.0 -> 0.52.4. Two specifiers had to change for the refresh to be honest: - litellm was capped at <1.85 when the provider was added (#1386, which records no reason for the bound), and dependabot later widened it to <1.86. That range ends inside a stretch of releases -- 1.83.8 through 1.92.x -- that declare Requires-Python >=3.10,<3.14. uv ignores that upper bound and installed 1.85.7 anyway, but pip does not: on Python 3.14, `pip install memmachine-server[litellm]` resolved to 1.83.7, not the 1.85.7 in the lock. 1.93.0 is the first release that supports 3.14, so that is the floor now, and the lock moves to 1.100.0. litellm also requires openai<3, which is what holds the OpenAI SDK on the 2.x line; without the floor the resolver satisfies that by walking litellm backwards to 1.83.0 instead. - boto3-stubs was pinned exactly at 1.43.13 while boto3 resolved to 1.43.90, so the stubs no longer described the installed SDK. Two source changes the new versions require, and one cleanup: - ty 0.0.79 rejects assigning over a bound method, so the metrics-factory wiring test installs its fake engine with monkeypatch.setattr. - `list(vector.flat)` is a list of numpy scalars, not `list[float]`, which the newer numpy typing now reports at the assignment. The replay path converts explicitly. - The dev group listed complexipy twice. Left behind deliberately: nebula5-python (see the cap and its reason), openai 3.x (litellm requires openai<3, and instructor's jiter<0.15 pin holds openai below 3.11 regardless; 3.0 also swaps the transport to httpx2), and ruff, still pinned at 0.15.14 -- 0.16 formats Markdown code blocks, which would reformat six docs, and adds LOG004, which flags four `logger.exception` calls, two of them genuine misuses and two in a helper that is only ever called from an except block. All six ty jobs pass, as do `ruff check`, `ruff format --check`, `uv lock --check` and the unit suite. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01BcPSoGsQjnVJS8A5NkoGZN * Drop dependencies nothing declares them for Four entries in the server's dependency list are not requirements of this package: - boto3-stubs is a type-stub distribution sitting in the runtime list. Nothing imports it, and `ty check --project packages/server` reports the same "All checks passed" with the package uninstalled, so it was buying neither runtime behaviour nor type coverage. (Base boto3-stubs only types the boto3 entry points; per-service clients need the mypy-boto3-* extras. If stronger boto3 typing is wanted later, that belongs in the dev group with those extras.) - regex belongs to memmachine-common, which declares it and imports it in api/spec.py. Nothing in the server package uses it. - The PyPI "dotenv" distribution ships no modules at all -- only dist-info -- and exists to depend on python-dotenv. `importlib.metadata.packages_distributions()["dotenv"]` is `["python-dotenv"]`, so `from dotenv import load_dotenv` has always been python-dotenv. Declare that instead. - greenlet is what SQLAlchemy's asyncio extra installs, and this package uses create_async_engine, so ask for `sqlalchemy[asyncio]` and let it bring greenlet. That also covers the platforms SQLAlchemy's own platform_machine marker on greenlet leaves out. The lock loses boto3-stubs, botocore-stubs, types-s3transfer and dotenv. greenlet and python-dotenv stay, now for stated reasons. reranker_manager took `runtime_checkable` from typing_extensions, an undeclared dependency it reached through pydantic. It has been in `typing` since 3.8 and this package requires 3.12, so it comes from `typing` now, next to the `Protocol` the file already imports there. memmachine-common and memmachine-client keep their typing_extensions dependency: they support 3.10, and `Self` and `Unpack` are 3.11. Deliberately kept, though nothing imports them, because each is named in the source or by the tests: - asyncpg and aiosqlite are the SQLAlchemy drivers this package builds URLs for. DatabaseBackend spells them out -- POSTGRES is ("postgres", SqlAlchemyConf, "postgresql", "asyncpg") -- SqlAlchemyConf renders `{dialect}+{driver}://`, and database_manager hands that to create_async_engine (and gates statement/connect timeouts on `conf.driver == "asyncpg"`). - tzdata on Windows. prompt_utilities builds `zoneinfo.ZoneInfo(tz)`, and Windows ships no system tz database, so without it every lookup -- including the UTC fallback in the except branch -- raises ZoneInfoNotFoundError. - milvus-lite, which pymilvus does not declare. It is what serves a file-backed Milvus URI, both for the local `path` deployment mode and for test_milvus_vector_store, which opens `MilvusClient(uri=tmp_path / "test_milvus.db")` behind `pytest.importorskip("milvus_lite")` -- 22 tests in the default run. All six ty jobs pass, as do `ruff check`, `ruff format --check`, `uv lock --check` and the unit suite. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01BcPSoGsQjnVJS8A5NkoGZN --------- Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
The lock had drifted far behind: 30 of the 52 direct dependencies were behind their latest release, some by a major version. Refreshing it moves 112 locked packages, adds 6 and removes 2, with no downgrades. Nine are major bumps: complexipy 6 -> 8, cryptography 49 -> 50, fastmcp 3 -> 4 (and fastmcp-slim), mcp 1 -> 2, sentence-transformers 5 -> 6, setuptools 83 -> 84, websockets 16 -> 17, xxhash 3 -> 4. Direct dependencies that moved include fastapi 0.139.0 -> 0.141.1, instructor 1.15.4 -> 1.17.0, langchain-aws 1.6.2 -> 1.7.6, neo4j 6.2.0 -> 6.3.1, openai 2.45.0 -> 2.54.0, qdrant-client 1.18.0 -> 1.19.0, sqlalchemy 2.0.51 -> 2.0.53, ty 0.0.59 -> 0.0.81 and uvicorn 0.51.0 -> 0.53.0. Two specifiers had to change for the refresh to be honest: - litellm was capped at <1.85 when the provider was added (MemMachine#1386, which records no reason for the bound), and dependabot later widened it to <1.86. That range ends inside a stretch of releases -- 1.83.8 through 1.92.x -- that declare Requires-Python <3.14. uv ignores that upper bound and installed 1.85.7 anyway, but pip does not: on Python 3.14, `pip install memmachine-server[litellm]` resolves to 1.83.7, not the 1.85.7 in the lock. 1.93.0 is the first release that supports 3.14, so that is the floor now, and the lock moves to 1.101.0. litellm also requires openai<3, which is what holds the OpenAI SDK on the 2.x line; without the floor the resolver satisfies that by walking litellm backwards to 1.83.0 instead. - boto3-stubs was pinned exactly at 1.43.13 while boto3 resolved to 1.43.47, so the stubs no longer described the installed SDK. One source change the new versions require, and one cleanup: - `list(vector.flat)` is a list of numpy scalars, not `list[float]`, which the newer numpy typing now reports at the assignment. The replay path converts explicitly. - The dev group listed complexipy twice. Left behind deliberately: nebula5-python (see the cap and its reason), openai 3.x (litellm requires openai<3, and instructor's jiter<0.15 pin holds openai below 3.11 regardless; 3.0 also swaps the transport to httpx2), and ruff, still pinned at 0.15.14 -- 0.16 formats Markdown code blocks, which would reformat six docs, and adds LOG004, which flags four `logger.exception` calls, two of them genuine misuses and two in a helper that is only ever called from an except block. Same change as the first half of MemMachine#1596 on speedkick, applied to main. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
…e ruff to 0.16 (#1646) * Bring the dependency lock up to date The lock had drifted far behind: 30 of the 52 direct dependencies were behind their latest release, some by a major version. Refreshing it moves 112 locked packages, adds 6 and removes 2, with no downgrades. Nine are major bumps: complexipy 6 -> 8, cryptography 49 -> 50, fastmcp 3 -> 4 (and fastmcp-slim), mcp 1 -> 2, sentence-transformers 5 -> 6, setuptools 83 -> 84, websockets 16 -> 17, xxhash 3 -> 4. Direct dependencies that moved include fastapi 0.139.0 -> 0.141.1, instructor 1.15.4 -> 1.17.0, langchain-aws 1.6.2 -> 1.7.6, neo4j 6.2.0 -> 6.3.1, openai 2.45.0 -> 2.54.0, qdrant-client 1.18.0 -> 1.19.0, sqlalchemy 2.0.51 -> 2.0.53, ty 0.0.59 -> 0.0.81 and uvicorn 0.51.0 -> 0.53.0. Two specifiers had to change for the refresh to be honest: - litellm was capped at <1.85 when the provider was added (#1386, which records no reason for the bound), and dependabot later widened it to <1.86. That range ends inside a stretch of releases -- 1.83.8 through 1.92.x -- that declare Requires-Python <3.14. uv ignores that upper bound and installed 1.85.7 anyway, but pip does not: on Python 3.14, `pip install memmachine-server[litellm]` resolves to 1.83.7, not the 1.85.7 in the lock. 1.93.0 is the first release that supports 3.14, so that is the floor now, and the lock moves to 1.101.0. litellm also requires openai<3, which is what holds the OpenAI SDK on the 2.x line; without the floor the resolver satisfies that by walking litellm backwards to 1.83.0 instead. - boto3-stubs was pinned exactly at 1.43.13 while boto3 resolved to 1.43.47, so the stubs no longer described the installed SDK. One source change the new versions require, and one cleanup: - `list(vector.flat)` is a list of numpy scalars, not `list[float]`, which the newer numpy typing now reports at the assignment. The replay path converts explicitly. - The dev group listed complexipy twice. Left behind deliberately: nebula5-python (see the cap and its reason), openai 3.x (litellm requires openai<3, and instructor's jiter<0.15 pin holds openai below 3.11 regardless; 3.0 also swaps the transport to httpx2), and ruff, still pinned at 0.15.14 -- 0.16 formats Markdown code blocks, which would reformat six docs, and adds LOG004, which flags four `logger.exception` calls, two of them genuine misuses and two in a helper that is only ever called from an except block. Same change as the first half of #1596 on speedkick, applied to main. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> * Drop dependencies nothing declares them for Four entries in the server's dependency list are not requirements of this package: - boto3-stubs is a type-stub distribution sitting in the runtime list. Nothing imports it, and `ty check --project packages/server` reports the same "All checks passed" with the package uninstalled, so it was buying neither runtime behavior nor type coverage. (Base boto3-stubs only types the boto3 entry points; per-service clients need the mypy-boto3-* extras. If stronger boto3 typing is wanted later, that belongs in the dev group with those extras.) - regex belongs to memmachine-common, which declares it and imports it in api/spec.py. Nothing in the server package uses it. - The PyPI "dotenv" distribution ships no modules at all -- only dist-info -- and exists to depend on python-dotenv. `importlib.metadata.packages_distributions()["dotenv"]` is `["python-dotenv"]`, so `from dotenv import load_dotenv` has always been python-dotenv. Declare that instead. - greenlet is what SQLAlchemy's asyncio extra installs, and this package uses create_async_engine, so ask for `sqlalchemy[asyncio]` and let it bring greenlet. That also covers the platforms SQLAlchemy's own platform_machine marker on greenlet leaves out. The lock loses boto3-stubs, botocore-stubs, types-s3transfer and dotenv. reranker_manager also took runtime_checkable from typing_extensions, an undeclared dependency it reached through pydantic. It has been in typing since 3.8 and this package requires 3.12, so it now comes from typing, next to the Protocol the file already imports there. memmachine-common and memmachine-client keep their typing_extensions dependency: they support 3.10, and Self and Unpack are 3.11. Same change as the second half of #1596 on speedkick, applied to main. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> * Log without a traceback where no exception is in flight openai_embedder's dimensionality check and amazon_bedrock_reranker's score-count check each call logger.exception() and then raise their own ExternalServiceAPIError. Nothing is being handled at that point: logger.exception is error(..., exc_info=True), and with no active exception it appends "NoneType: None" to the record. logger.error is what both sites mean; the raise that follows carries the message. ruff 0.16 reports both as LOG004 (`.exception()` call outside exception handlers). Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> * Re-raise the HTTPError the client's project-error helper is given _handle_get_project_http_error takes the caught HTTPError as a parameter, but its bare `raise` statements and logger.exception() calls ignored it and used the exception the *caller's* except block had active. That works only while the helper is called from inside that block, and ruff 0.16's LOG004 cannot see the call boundary, so it reports the two logging calls as .exception() outside a handler. The helper now uses what it is handed: logger.error(..., exc_info=error) attaches the same traceback logger.exception would have, and `raise error` re-raises the same object. Callers observe no difference: the exception identity, its original traceback and the 422 ValueError's __cause__ are unchanged (the helper's own frame is appended to the traceback, as it is for any re-raise from a callee). Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> * Move ruff to 0.16.8 ruff was held at 0.15.14 by the dependency refresh because 0.16 brings two things that touch committed content, and both now have their own commits ahead of this one: - LOG004 (`.exception()` outside an exception handler), selected via the LOG group. Its four findings are resolved by the two preceding commits, so `ruff check` is clean at this pin with no suppressions. - The formatter now formats Python code blocks inside Markdown. The six docs it touches are reformatted here, by `ruff format` and nothing else: USAGE.md, examples/v1/README.md, integrations/aws_strands_agent_sdk/README.md, integrations/langgraph/README.md, maintainers/build-pip-packages.md, packages/client/README.md. Trailing commas, collapsed argument lists, and column-aligned comments brought to two spaces; no words change. The lint workflow's ruff-action reads the pin from pyproject.toml, so CI enforces 0.16.8 from this commit on. No .py file is reformatted by the new version. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> * Explain only the litellm floor in the server's pyproject The litellm extra's comment also said litellm's openai<3 is what keeps the resolved OpenAI SDK on the 2.x line. That describes the universal lock, which resolves every extra, and not a requirement of this package: the server hands the SDK no httpx client, which is the only thing openai 3.0 changed, and its suite passes on openai 3.3.0 with the extra absent (1869 passed, 3 skipped, the same as on the lock). Installs without the extra may resolve openai 3.x. The comment now states only the floor, which is this package's own decision. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> --------- Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
Purpose of the change
Today every new LLM backing in MemMachine (Cohere, Mistral, Groq, Together, ...) requires writing another
LanguageModelfrom scratch alongsideOpenAIChatCompletionsLanguageModel/OpenAIResponsesLanguageModel/AmazonBedrockLanguageModel. This PR adds a singleLiteLLMLanguageModelthat delegates the actual provider call to the LiteLLM SDK, givingMemMachine coverage of the 100+ providers LiteLLM supports (OpenAI, Anthropic, AWS Bedrock, Vertex AI, Cohere, Mistral, Groq, Perplexity, Together, Fireworks, Cerebras, Databricks, IBM Watsonx, AI21,
Replicate, DeepInfra, NVIDIA NIM, xAI, Sambanova, ...) by changing only the
modelspec.It also adds a third deployment mode (LiteLLM proxy server) useful for centralized credential management and audit logging.
Description
LiteLLMLanguageModelsubclassesOpenAIChatCompletionsLanguageModeland overrides only_request_chat_completionto calllitellm.acompletion(**args)instead ofclient.chat.completions.create(**args).LiteLLM normalizes every backing's response to OpenAI's
ChatCompletionshape, so the parent's parsing, streaming, tool-call accumulation, structured-output handling, and metrics paths inherit unchanged.Configuration mirrors the existing OpenAI shape:
In embedded mode (no
api_base), LiteLLM resolves credentials from each backing's standard env var (ANTHROPIC_API_KEY,OPENAI_API_KEY,AWS_ACCESS_KEY_ID, ...) at call time. In proxy mode, calls routethrough a LiteLLM proxy server that holds the credentials.
Dependency:
litellm>=1.60,<1.85. Imported lazily inside the request function so users who don't configure alitellm_language_model_confsentry don't need it. Happy to make this an optional extra(
memmachine-server[litellm]) instead if preferred.Files added:
packages/server/src/memmachine_server/common/language_model/litellm_language_model.py(new, 224 LOC)packages/server/server_tests/memmachine_server/common/language_model/test_litellm_language_model.py(new, 270 LOC)Files modified:
packages/server/src/memmachine_server/common/configuration/language_model_conf.py(+71): newLiteLLMLanguageModelConf,litellm_language_model_confsdict onLanguageModelsConf, helper accessors.packages/server/src/memmachine_server/common/resource_manager/language_model_manager.py(+50): wiredlitellminto_is_configured,get_all_names,_build_language_model,add_language_model_config,remove_language_model; new_build_litellm_language_modelbuilder.Fixes/Closes
N/A (no related issue; happy to open one and link if preferred).
Type of change
How Has This Been Tested?
Unit tests: 10 new tests in
test_litellm_language_model.pycovering init, dispatch, kwarg forwarding, retry behavior, and inherited parsing.Coverage:
LiteLLMLanguageModel.__init__does not require anAsyncOpenAIclient (parent's hard requirement is bypassed cleanly)._request_chat_completiondispatches tolitellm.acompletionwith the right model spec.api_key/api_base/api_versionare forwarded only when set on params; caller-supplied kwargs win over shim defaults (no silent override).extra_kwargsfrom params (metadata,tags, ...) reachlitellm.acompletion.RateLimitError,APITimeoutError,APIConnectionError,InternalServerError,ServiceUnavailableError,Timeout) trigger exponential backoff; non-retryable ones raise immediately.max_attempts, retryable errors surface asExternalServiceAPIError.generate_responseinherits the parent's OpenAI-shape parsing unchanged when_request_chat_completionreturns aChatCompletion.Lint + type-check (CI parity):
End-to-end test (Anthropic via Azure AI Foundry):
Output:
'pong.'. The wrapped call routed throughlitellm.acompletionto Anthropic and the response was parsed by the inherited OpenAI parser. The sameLiteLLMLanguageModelwould route via OpenAI / Bedrock/ Cohere / Mistral / ... by changing only the
modelspec.Test Results: All 10 unit tests pass; lint, format, and ty are clean; live E2E returns the expected reply.
Checklist
-sSper CONTRIBUTING.md; will rebase before merge if needed)Maintainer Checklist
Screenshots/Gifs
N/A (backend-only change; live E2E output included in "How Has This Been Tested?").
Further comments
Out of scope (happy to follow up):
LiteLLMEmbedder. LiteLLM also exposeslitellm.aembedding(Cohere, Voyage, Mistral, Bedrock Titan, Vertex, ...). Glad to ship this in a separate PR if you'd like the same single-implementationcoverage on the embedder side.
litellmtodependenciesdirectly. Currently lazy-imported inside the request function; can promote to an optional extramemmachine-server[litellm]if you'd prefer it gated.