Skip to content

feat(language_model): add LiteLLM provider for 100+ backings - #1386

Merged
sscargal merged 3 commits into
MemMachine:mainfrom
RheagalFire:feat/add-litellm-language-model
May 19, 2026
Merged

sscargal merged 3 commits into
MemMachine:mainfrom
RheagalFire:feat/add-litellm-language-model

Conversation

@RheagalFire

Copy link
Copy Markdown
Contributor

Purpose of the change

Today every new LLM backing in MemMachine (Cohere, Mistral, Groq, Together, ...) requires writing another LanguageModel from scratch alongside OpenAIChatCompletionsLanguageModel /
OpenAIResponsesLanguageModel / AmazonBedrockLanguageModel. This PR adds a single LiteLLMLanguageModel that delegates the actual provider call to the LiteLLM SDK, giving
MemMachine coverage of the 100+ providers LiteLLM supports (OpenAI, Anthropic, AWS Bedrock, Vertex AI, Cohere, Mistral, Groq, Perplexity, Together, Fireworks, Cerebras, Databricks, IBM Watsonx, AI21,
Replicate, DeepInfra, NVIDIA NIM, xAI, Sambanova, ...) by changing only the model spec.

It also adds a third deployment mode (LiteLLM proxy server) useful for centralized credential management and audit logging.

Description

LiteLLMLanguageModel subclasses OpenAIChatCompletionsLanguageModel and overrides only _request_chat_completion to call litellm.acompletion(**args) instead of client.chat.completions.create(**args).
LiteLLM normalizes every backing's response to OpenAI's ChatCompletion shape, so the parent's parsing, streaming, tool-call accumulation, structured-output handling, and metrics paths inherit unchanged.

Configuration mirrors the existing OpenAI shape:

language_models:                                                                                                                                                                                                   
  litellm_language_model_confs:            
    sonnet:                                                                                                                                                                                                        
      model: anthropic/claude-sonnet-4-6                     
    gpt4o:                                                                                                                                                                                                         
      model: openai/gpt-4o                                                                                                                                                                                         
    cohere:                                                  
      model: cohere/command-r-plus-08-2024                                                                                                                                                                         
                                                  
    # Proxy mode (centralized credentials)                   
    proxied:                                                                                                                                                                                                       
      model: anthropic/claude-sonnet-4-6                     
      api_base: http://localhost:4000                                                                                                                                                                              
      api_key: sk-fastagent-proxy-1234                                                                                                                                                                             

In embedded mode (no api_base), LiteLLM resolves credentials from each backing's standard env var (ANTHROPIC_API_KEY, OPENAI_API_KEY, AWS_ACCESS_KEY_ID, ...) at call time. In proxy mode, calls route
through a LiteLLM proxy server that holds the credentials.

Dependency: litellm>=1.60,<1.85. Imported lazily inside the request function so users who don't configure a litellm_language_model_confs entry don't need it. Happy to make this an optional extra
(memmachine-server[litellm]) instead if preferred.

Files added:

  • packages/server/src/memmachine_server/common/language_model/litellm_language_model.py (new, 224 LOC)
  • packages/server/server_tests/memmachine_server/common/language_model/test_litellm_language_model.py (new, 270 LOC)

Files modified:

  • packages/server/src/memmachine_server/common/configuration/language_model_conf.py (+71): new LiteLLMLanguageModelConf, litellm_language_model_confs dict on LanguageModelsConf, helper accessors.
  • packages/server/src/memmachine_server/common/resource_manager/language_model_manager.py (+50): wired litellm into _is_configured, get_all_names, _build_language_model, add_language_model_config,
    remove_language_model; new _build_litellm_language_model builder.

Fixes/Closes

N/A (no related issue; happy to open one and link if preferred).

Type of change

  • New feature (non-breaking change which adds functionality)

How Has This Been Tested?

  • Unit Test
  • End-to-end Test (manual against a real provider)

Unit tests: 10 new tests in test_litellm_language_model.py covering init, dispatch, kwarg forwarding, retry behavior, and inherited parsing.

$ pytest packages/server/server_tests/memmachine_server/common/language_model/test_litellm_language_model.py -v                                                                                                    
                                                  
test_init_does_not_require_openai_client                                       PASSED                                                                                                                              
test_request_chat_completion_dispatches_to_litellm                             PASSED
test_request_chat_completion_forwards_api_base_and_key                         PASSED                                                                                                                              
test_request_chat_completion_does_not_overwrite_explicit_kwargs                PASSED                                                                                                                              
test_request_chat_completion_extra_kwargs_forwarded                            PASSED                                                                                                                              
test_request_chat_completion_retries_on_retryable_error                        PASSED                                                                                                                              
test_request_chat_completion_raises_external_service_error_after_max_attempts  PASSED
test_request_chat_completion_non_retryable_error_raises_immediately            PASSED                                                                                                                              
test_is_retryable_litellm_error_recognizes_known_classes                       PASSED
test_generate_response_uses_parent_parsing                                     PASSED                                                                                                                              
                                                                                                                                                                                                                   
10 passed in 5.24s                                                                                                                                                                                                 

Coverage:

  • LiteLLMLanguageModel.__init__ does not require an AsyncOpenAI client (parent's hard requirement is bypassed cleanly).
  • _request_chat_completion dispatches to litellm.acompletion with the right model spec.
  • api_key / api_base / api_version are forwarded only when set on params; caller-supplied kwargs win over shim defaults (no silent override).
  • extra_kwargs from params (metadata, tags, ...) reach litellm.acompletion.
  • Retryable errors (RateLimitError, APITimeoutError, APIConnectionError, InternalServerError, ServiceUnavailableError, Timeout) trigger exponential backoff; non-retryable ones raise immediately.
  • After max_attempts, retryable errors surface as ExternalServiceAPIError.
  • generate_response inherits the parent's OpenAI-shape parsing unchanged when _request_chat_completion returns a ChatCompletion.

Lint + type-check (CI parity):

$ ruff check <touched files>           ─►  All checks passed!                                                                                                                                                      
$ ruff format --check <touched files>  ─►  3 files already formatted                                                                                                                                               
$ ty check --project packages/server --python-version 3.12 <touched files>  ─►  All checks passed!                                                                                                                 

End-to-end test (Anthropic via Azure AI Foundry):

import asyncio                                                                                                                                                                                                     
from memmachine_server.common.language_model.litellm_language_model import (                                                                                                                                       
    LiteLLMLanguageModel, LiteLLMLanguageModelParams,                                                                                                                                                              
)                                                                                                                                                                                                                  
                                                                                                                                                                                                                   
async def main():                                                                                                                                                                                                  
    lm = LiteLLMLanguageModel(LiteLLMLanguageModelParams(                                                                                                                                                          
        model="anthropic/claude-sonnet-4-6",                                                                                                                                                                       
    ))                                            
    text, tools = await lm.generate_response(                                                                                                                                                                      
        system_prompt="You answer with a single word.",                                                                                                                                                          
        user_prompt="Reply with: pong.",                                                                                                                                                                           
    )                                                                                                                                                                                                            
    print(text)        # 'pong.'                                                                                                                                                                                   
    print(tools)       # []                                   
                                                                                                                                                                                                                   
asyncio.run(main())                               

Output: 'pong.'. The wrapped call routed through litellm.acompletion to Anthropic and the response was parsed by the inherited OpenAI parser. The same LiteLLMLanguageModel would route via OpenAI / Bedrock
/ Cohere / Mistral / ... by changing only the model spec.

Test Results: All 10 unit tests pass; lint, format, and ty are clean; live E2E returns the expected reply.

Checklist

  • I have signed the commit(s) within this pull request (needs -sS per CONTRIBUTING.md; will rebase before merge if needed)
  • My code follows the style guidelines of this project (See STYLE_GUIDE.md)
  • I have performed a self-review of my own code
  • I have commented my code
  • I have made corresponding changes to the documentation (N/A; new optional provider; happy to add docs in a follow-up)
  • My changes generate no new warnings
  • I have added unit tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes
  • Any dependent changes have been merged and published in downstream modules
  • I have checked my code and corrected any misspellings

Maintainer Checklist

  • Confirmed all checks passed
  • Contributor has signed the commit(s)
  • Reviewed the code
  • Run, Tested, and Verified the change(s) work as expected

Screenshots/Gifs

N/A (backend-only change; live E2E output included in "How Has This Been Tested?").

Further comments

Out of scope (happy to follow up):

  • A parallel LiteLLMEmbedder. LiteLLM also exposes litellm.aembedding (Cohere, Voyage, Mistral, Bedrock Titan, Vertex, ...). Glad to ship this in a separate PR if you'd like the same single-implementation
    coverage on the embedder side.
  • Adding litellm to dependencies directly. Currently lazy-imported inside the request function; can promote to an optional extra memmachine-server[litellm] if you'd prefer it gated.

@RheagalFire

Copy link
Copy Markdown
Contributor Author

cc @sscargal

@malatewang
malatewang requested review from edwinyyyu, jealous and malatewang and removed request for edwinyyyu and jealous May 1, 2026 21:40

@edwinyyyu edwinyyyu left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a reason not to add litellm as an optional dependency in server pyproject.toml?

def get_litellm_language_model_conf(self, name: str) -> "LiteLLMLanguageModelConf":
"""Get LiteLLM language model configuration by name."""
return self.litellm_language_model_confs[name]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The parse and to_yaml_dict needs update too

max_attempts: int,
generate_response_call_uuid: object,
) -> ChatCompletion | AsyncIterator[object] | object:
import litellm

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

update the project.toml file to install the module

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds a new server-side LiteLLMLanguageModel provider intended to expand MemMachine’s supported LLM backends by delegating requests to the LiteLLM SDK while reusing the existing OpenAI chat-completions parsing/streaming/tool-call logic.

Changes:

  • Introduces LiteLLMLanguageModel (subclassing OpenAIChatCompletionsLanguageModel) that routes requests via litellm.acompletion.
  • Extends language model configuration and the LanguageModelManager to register/build/remove LiteLLM-backed models.
  • Adds unit tests for the LiteLLM adapter behavior (dispatch, kwarg forwarding, retries, and parent parsing).

Reviewed changes

Copilot reviewed 4 out of 4 changed files in this pull request and generated 3 comments.

File Description
packages/server/src/memmachine_server/common/resource_manager/language_model_manager.py Wires LiteLLM configs into manager lookups and adds a builder for LiteLLMLanguageModel.
packages/server/src/memmachine_server/common/language_model/litellm_language_model.py New LiteLLM adapter that swaps the request implementation to litellm.acompletion and adds retry logic.
packages/server/src/memmachine_server/common/configuration/language_model_conf.py Adds LiteLLMLanguageModelConf and a litellm_language_model_confs collection plus accessors.
packages/server/server_tests/memmachine_server/common/language_model/test_litellm_language_model.py New unit tests validating LiteLLM dispatch/forwarding/retry behavior and inherited parsing.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +10 to +12
from typing import Any
from unittest.mock import AsyncMock, MagicMock, patch

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 5 out of 5 changed files in this pull request and generated 4 comments.

Comment on lines 180 to +188
ret: LanguageModel | None = None
if name in self.conf.openai_responses_language_model_confs:
ret = self._build_openai_responses_language_model(name)
if name in self.conf.openai_chat_completions_language_model_confs:
ret = self._build_openai_chat_completions_language_model(name)
if name in self.conf.amazon_bedrock_language_model_confs:
ret = self._build_amazon_bedrock_language_model(name)
if name in self.conf.litellm_language_model_confs:
ret = self._build_litellm_language_model(name)
Comment on lines +63 to +65
litellm = [
"litellm>=1.63.0",
]
Comment on lines +159 to +165
try:
import litellm
except ImportError as e:
raise ImportError(
"litellm is required for LiteLLMLanguageModel. "
"Install it with: pip install memmachine-server[litellm]"
) from e
@edwinyyyu

Copy link
Copy Markdown
Contributor

Please ensure CI passes.

@sscargal sscargal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@RheagalFire, in addition to @edwinyyyu and CoPilot feedback, please sign your commits. Thanks.

@RheagalFire
RheagalFire force-pushed the feat/add-litellm-language-model branch 3 times, most recently from 711c275 to 6b5e6ee Compare May 11, 2026 22:21
@RheagalFire
RheagalFire force-pushed the feat/add-litellm-language-model branch from c34dccb to d03abf1 Compare May 18, 2026 21:29
@RheagalFire

Copy link
Copy Markdown
Contributor Author

@sscargal @edwinyyyu

Signed the commits. All previous feedback (parse/to_yaml_dict, optional dep in server pyproject.toml, lint/type check fixes) was already addressed in earlier commits -- now consolidated into one clean commit.

@sscargal sscargal added this to the v0.4.0 milestone May 19, 2026
@sscargal

Copy link
Copy Markdown
Contributor

Many thanks @RheagalFire . There's a format --check --diff failure that's a quick fix. If you can implement it, then it's ready to merge.

@RheagalFire

Copy link
Copy Markdown
Contributor Author

@sscargal can you approve the CI.

@sscargal sscargal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks for the new feature.

@sscargal
sscargal merged commit 1b1dbac into MemMachine:main May 19, 2026
44 of 46 checks passed
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 9, 2026
The lock had drifted far behind: 30 of the 52 direct dependencies were
behind their latest release, some by a major version. Refreshing it moves
104 locked packages, adds 6 and removes 2, with no downgrades. Nine are
major bumps: complexipy 6 -> 8, cryptography 49 -> 50, fastmcp 3 -> 4 (and
fastmcp-slim), mcp 1 -> 2, sentence-transformers 5 -> 6, setuptools 83 ->
84, websockets 16 -> 17, xxhash 3 -> 4. Direct dependencies that moved
include fastapi 0.139.0 -> 0.141.1, instructor 1.15.4 -> 1.17.0, langchain-aws
1.6.2 -> 1.7.5, neo4j 6.2.0 -> 6.3.0, openai 2.45.0 -> 2.54.0, qdrant-client
1.18.0 -> 1.19.0, sqlalchemy 2.0.51 -> 2.0.52, ty 0.0.59 -> 0.0.79 and
uvicorn 0.51.0 -> 0.52.4.

Two specifiers had to change for the refresh to be honest:

- litellm was capped at <1.85 when the provider was added (MemMachine#1386, which
  records no reason for the bound), and dependabot later widened it to
  <1.86. That range ends inside a stretch of releases -- 1.83.8 through
  1.92.x -- that declare Requires-Python >=3.10,<3.14. uv ignores that
  upper bound and installed 1.85.7 anyway, but pip does not: on Python
  3.14, `pip install memmachine-server[litellm]` resolved to 1.83.7, not
  the 1.85.7 in the lock. 1.93.0 is the first release that supports 3.14,
  so that is the floor now, and the lock moves to 1.100.0. litellm also
  requires openai<3, which is what holds the OpenAI SDK on the 2.x line;
  without the floor the resolver satisfies that by walking litellm
  backwards to 1.83.0 instead.
- boto3-stubs was pinned exactly at 1.43.13 while boto3 resolved to
  1.43.90, so the stubs no longer described the installed SDK.

Two source changes the new versions require, and one cleanup:

- ty 0.0.79 rejects assigning over a bound method, so the metrics-factory
  wiring test installs its fake engine with monkeypatch.setattr.
- `list(vector.flat)` is a list of numpy scalars, not `list[float]`, which
  the newer numpy typing now reports at the assignment. The replay path
  converts explicitly.
- The dev group listed complexipy twice.

Left behind deliberately: nebula5-python (see the cap and its reason),
openai 3.x (litellm requires openai<3, and instructor's jiter<0.15 pin
holds openai below 3.11 regardless; 3.0 also swaps the transport to
httpx2), and ruff, still pinned at 0.15.14 -- 0.16 formats Markdown code
blocks, which would reformat six docs, and adds LOG004, which flags four
`logger.exception` calls, two of them genuine misuses and two in a helper
that is only ever called from an except block.

All six ty jobs pass, as do `ruff check`, `ruff format --check`,
`uv lock --check` and the unit suite.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BcPSoGsQjnVJS8A5NkoGZN
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 9, 2026
The lock had drifted far behind: 30 of the 52 direct dependencies were
behind their latest release, some by a major version. Refreshing it moves
104 locked packages, adds 6 and removes 2, with no downgrades. Nine are
major bumps: complexipy 6 -> 8, cryptography 49 -> 50, fastmcp 3 -> 4 (and
fastmcp-slim), mcp 1 -> 2, sentence-transformers 5 -> 6, setuptools 83 ->
84, websockets 16 -> 17, xxhash 3 -> 4. Direct dependencies that moved
include fastapi 0.139.0 -> 0.141.1, instructor 1.15.4 -> 1.17.0, langchain-aws
1.6.2 -> 1.7.5, neo4j 6.2.0 -> 6.3.0, openai 2.45.0 -> 2.54.0, qdrant-client
1.18.0 -> 1.19.0, sqlalchemy 2.0.51 -> 2.0.52, ty 0.0.59 -> 0.0.79 and
uvicorn 0.51.0 -> 0.52.4.

Two specifiers had to change for the refresh to be honest:

- litellm was capped at <1.85 when the provider was added (MemMachine#1386, which
  records no reason for the bound), and dependabot later widened it to
  <1.86. That range ends inside a stretch of releases -- 1.83.8 through
  1.92.x -- that declare Requires-Python >=3.10,<3.14. uv ignores that
  upper bound and installed 1.85.7 anyway, but pip does not: on Python
  3.14, `pip install memmachine-server[litellm]` resolved to 1.83.7, not
  the 1.85.7 in the lock. 1.93.0 is the first release that supports 3.14,
  so that is the floor now, and the lock moves to 1.100.0. litellm also
  requires openai<3, which is what holds the OpenAI SDK on the 2.x line;
  without the floor the resolver satisfies that by walking litellm
  backwards to 1.83.0 instead.
- boto3-stubs was pinned exactly at 1.43.13 while boto3 resolved to
  1.43.90, so the stubs no longer described the installed SDK.

Two source changes the new versions require, and one cleanup:

- ty 0.0.79 rejects assigning over a bound method, so the metrics-factory
  wiring test installs its fake engine with monkeypatch.setattr.
- `list(vector.flat)` is a list of numpy scalars, not `list[float]`, which
  the newer numpy typing now reports at the assignment. The replay path
  converts explicitly.
- The dev group listed complexipy twice.

Left behind deliberately: nebula5-python (see the cap and its reason),
openai 3.x (litellm requires openai<3, and instructor's jiter<0.15 pin
holds openai below 3.11 regardless; 3.0 also swaps the transport to
httpx2), and ruff, still pinned at 0.15.14 -- 0.16 formats Markdown code
blocks, which would reformat six docs, and adds LOG004, which flags four
`logger.exception` calls, two of them genuine misuses and two in a helper
that is only ever called from an except block.

All six ty jobs pass, as do `ruff check`, `ruff format --check`,
`uv lock --check` and the unit suite.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BcPSoGsQjnVJS8A5NkoGZN
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 10, 2026
The lock had drifted far behind: 30 of the 52 direct dependencies were
behind their latest release, some by a major version. Refreshing it moves
104 locked packages, adds 6 and removes 2, with no downgrades. Nine are
major bumps: complexipy 6 -> 8, cryptography 49 -> 50, fastmcp 3 -> 4 (and
fastmcp-slim), mcp 1 -> 2, sentence-transformers 5 -> 6, setuptools 83 ->
84, websockets 16 -> 17, xxhash 3 -> 4. Direct dependencies that moved
include fastapi 0.139.0 -> 0.141.1, instructor 1.15.4 -> 1.17.0, langchain-aws
1.6.2 -> 1.7.5, neo4j 6.2.0 -> 6.3.0, openai 2.45.0 -> 2.54.0, qdrant-client
1.18.0 -> 1.19.0, sqlalchemy 2.0.51 -> 2.0.52, ty 0.0.59 -> 0.0.79 and
uvicorn 0.51.0 -> 0.52.4.

Two specifiers had to change for the refresh to be honest:

- litellm was capped at <1.85 when the provider was added (MemMachine#1386, which
  records no reason for the bound), and dependabot later widened it to
  <1.86. That range ends inside a stretch of releases -- 1.83.8 through
  1.92.x -- that declare Requires-Python >=3.10,<3.14. uv ignores that
  upper bound and installed 1.85.7 anyway, but pip does not: on Python
  3.14, `pip install memmachine-server[litellm]` resolved to 1.83.7, not
  the 1.85.7 in the lock. 1.93.0 is the first release that supports 3.14,
  so that is the floor now, and the lock moves to 1.100.0. litellm also
  requires openai<3, which is what holds the OpenAI SDK on the 2.x line;
  without the floor the resolver satisfies that by walking litellm
  backwards to 1.83.0 instead.
- boto3-stubs was pinned exactly at 1.43.13 while boto3 resolved to
  1.43.90, so the stubs no longer described the installed SDK.

Two source changes the new versions require, and one cleanup:

- ty 0.0.79 rejects assigning over a bound method, so the metrics-factory
  wiring test installs its fake engine with monkeypatch.setattr.
- `list(vector.flat)` is a list of numpy scalars, not `list[float]`, which
  the newer numpy typing now reports at the assignment. The replay path
  converts explicitly.
- The dev group listed complexipy twice.

Left behind deliberately: nebula5-python (see the cap and its reason),
openai 3.x (litellm requires openai<3, and instructor's jiter<0.15 pin
holds openai below 3.11 regardless; 3.0 also swaps the transport to
httpx2), and ruff, still pinned at 0.15.14 -- 0.16 formats Markdown code
blocks, which would reformat six docs, and adds LOG004, which flags four
`logger.exception` calls, two of them genuine misuses and two in a helper
that is only ever called from an except block.

All six ty jobs pass, as do `ruff check`, `ruff format --check`,
`uv lock --check` and the unit suite.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BcPSoGsQjnVJS8A5NkoGZN
edwinyyyu added a commit that referenced this pull request Sep 10, 2026
…edkick) (#1596)

* Bring the dependency lock up to date

The lock had drifted far behind: 30 of the 52 direct dependencies were
behind their latest release, some by a major version. Refreshing it moves
104 locked packages, adds 6 and removes 2, with no downgrades. Nine are
major bumps: complexipy 6 -> 8, cryptography 49 -> 50, fastmcp 3 -> 4 (and
fastmcp-slim), mcp 1 -> 2, sentence-transformers 5 -> 6, setuptools 83 ->
84, websockets 16 -> 17, xxhash 3 -> 4. Direct dependencies that moved
include fastapi 0.139.0 -> 0.141.1, instructor 1.15.4 -> 1.17.0, langchain-aws
1.6.2 -> 1.7.5, neo4j 6.2.0 -> 6.3.0, openai 2.45.0 -> 2.54.0, qdrant-client
1.18.0 -> 1.19.0, sqlalchemy 2.0.51 -> 2.0.52, ty 0.0.59 -> 0.0.79 and
uvicorn 0.51.0 -> 0.52.4.

Two specifiers had to change for the refresh to be honest:

- litellm was capped at <1.85 when the provider was added (#1386, which
  records no reason for the bound), and dependabot later widened it to
  <1.86. That range ends inside a stretch of releases -- 1.83.8 through
  1.92.x -- that declare Requires-Python >=3.10,<3.14. uv ignores that
  upper bound and installed 1.85.7 anyway, but pip does not: on Python
  3.14, `pip install memmachine-server[litellm]` resolved to 1.83.7, not
  the 1.85.7 in the lock. 1.93.0 is the first release that supports 3.14,
  so that is the floor now, and the lock moves to 1.100.0. litellm also
  requires openai<3, which is what holds the OpenAI SDK on the 2.x line;
  without the floor the resolver satisfies that by walking litellm
  backwards to 1.83.0 instead.
- boto3-stubs was pinned exactly at 1.43.13 while boto3 resolved to
  1.43.90, so the stubs no longer described the installed SDK.

Two source changes the new versions require, and one cleanup:

- ty 0.0.79 rejects assigning over a bound method, so the metrics-factory
  wiring test installs its fake engine with monkeypatch.setattr.
- `list(vector.flat)` is a list of numpy scalars, not `list[float]`, which
  the newer numpy typing now reports at the assignment. The replay path
  converts explicitly.
- The dev group listed complexipy twice.

Left behind deliberately: nebula5-python (see the cap and its reason),
openai 3.x (litellm requires openai<3, and instructor's jiter<0.15 pin
holds openai below 3.11 regardless; 3.0 also swaps the transport to
httpx2), and ruff, still pinned at 0.15.14 -- 0.16 formats Markdown code
blocks, which would reformat six docs, and adds LOG004, which flags four
`logger.exception` calls, two of them genuine misuses and two in a helper
that is only ever called from an except block.

All six ty jobs pass, as do `ruff check`, `ruff format --check`,
`uv lock --check` and the unit suite.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BcPSoGsQjnVJS8A5NkoGZN

* Drop dependencies nothing declares them for

Four entries in the server's dependency list are not requirements of this
package:

- boto3-stubs is a type-stub distribution sitting in the runtime list.
  Nothing imports it, and `ty check --project packages/server` reports
  the same "All checks passed" with the package uninstalled, so it was
  buying neither runtime behaviour nor type coverage. (Base boto3-stubs
  only types the boto3 entry points; per-service clients need the
  mypy-boto3-* extras. If stronger boto3 typing is wanted later, that
  belongs in the dev group with those extras.)
- regex belongs to memmachine-common, which declares it and imports it in
  api/spec.py. Nothing in the server package uses it.
- The PyPI "dotenv" distribution ships no modules at all -- only
  dist-info -- and exists to depend on python-dotenv.
  `importlib.metadata.packages_distributions()["dotenv"]` is
  `["python-dotenv"]`, so `from dotenv import load_dotenv` has always
  been python-dotenv. Declare that instead.
- greenlet is what SQLAlchemy's asyncio extra installs, and this package
  uses create_async_engine, so ask for `sqlalchemy[asyncio]` and let it
  bring greenlet. That also covers the platforms SQLAlchemy's own
  platform_machine marker on greenlet leaves out.

The lock loses boto3-stubs, botocore-stubs, types-s3transfer and dotenv.
greenlet and python-dotenv stay, now for stated reasons.

reranker_manager took `runtime_checkable` from typing_extensions, an
undeclared dependency it reached through pydantic. It has been in
`typing` since 3.8 and this package requires 3.12, so it comes from
`typing` now, next to the `Protocol` the file already imports there.
memmachine-common and memmachine-client keep their typing_extensions
dependency: they support 3.10, and `Self` and `Unpack` are 3.11.

Deliberately kept, though nothing imports them, because each is named in
the source or by the tests:

- asyncpg and aiosqlite are the SQLAlchemy drivers this package builds
  URLs for. DatabaseBackend spells them out -- POSTGRES is
  ("postgres", SqlAlchemyConf, "postgresql", "asyncpg") -- SqlAlchemyConf
  renders `{dialect}+{driver}://`, and database_manager hands that to
  create_async_engine (and gates statement/connect timeouts on
  `conf.driver == "asyncpg"`).
- tzdata on Windows. prompt_utilities builds `zoneinfo.ZoneInfo(tz)`, and
  Windows ships no system tz database, so without it every lookup --
  including the UTC fallback in the except branch -- raises
  ZoneInfoNotFoundError.
- milvus-lite, which pymilvus does not declare. It is what serves a
  file-backed Milvus URI, both for the local `path` deployment mode and
  for test_milvus_vector_store, which opens
  `MilvusClient(uri=tmp_path / "test_milvus.db")` behind
  `pytest.importorskip("milvus_lite")` -- 22 tests in the default run.

All six ty jobs pass, as do `ruff check`, `ruff format --check`,
`uv lock --check` and the unit suite.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01BcPSoGsQjnVJS8A5NkoGZN

---------

Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 16, 2026
The lock had drifted far behind: 30 of the 52 direct dependencies were
behind their latest release, some by a major version. Refreshing it moves
112 locked packages, adds 6 and removes 2, with no downgrades. Nine are
major bumps: complexipy 6 -> 8, cryptography 49 -> 50, fastmcp 3 -> 4 (and
fastmcp-slim), mcp 1 -> 2, sentence-transformers 5 -> 6, setuptools 83 ->
84, websockets 16 -> 17, xxhash 3 -> 4. Direct dependencies that moved
include fastapi 0.139.0 -> 0.141.1, instructor 1.15.4 -> 1.17.0, langchain-aws
1.6.2 -> 1.7.6, neo4j 6.2.0 -> 6.3.1, openai 2.45.0 -> 2.54.0, qdrant-client
1.18.0 -> 1.19.0, sqlalchemy 2.0.51 -> 2.0.53, ty 0.0.59 -> 0.0.81 and
uvicorn 0.51.0 -> 0.53.0.

Two specifiers had to change for the refresh to be honest:

- litellm was capped at <1.85 when the provider was added (MemMachine#1386, which
  records no reason for the bound), and dependabot later widened it to
  <1.86. That range ends inside a stretch of releases -- 1.83.8 through
  1.92.x -- that declare Requires-Python <3.14. uv ignores that upper
  bound and installed 1.85.7 anyway, but pip does not: on Python 3.14,
  `pip install memmachine-server[litellm]` resolves to 1.83.7, not the
  1.85.7 in the lock. 1.93.0 is the first release that supports 3.14,
  so that is the floor now, and the lock moves to 1.101.0. litellm also
  requires openai<3, which is what holds the OpenAI SDK on the 2.x line;
  without the floor the resolver satisfies that by walking litellm
  backwards to 1.83.0 instead.
- boto3-stubs was pinned exactly at 1.43.13 while boto3 resolved to
  1.43.47, so the stubs no longer described the installed SDK.

One source change the new versions require, and one cleanup:

- `list(vector.flat)` is a list of numpy scalars, not `list[float]`, which
  the newer numpy typing now reports at the assignment. The replay path
  converts explicitly.
- The dev group listed complexipy twice.

Left behind deliberately: nebula5-python (see the cap and its reason),
openai 3.x (litellm requires openai<3, and instructor's jiter<0.15 pin
holds openai below 3.11 regardless; 3.0 also swaps the transport to
httpx2), and ruff, still pinned at 0.15.14 -- 0.16 formats Markdown code
blocks, which would reformat six docs, and adds LOG004, which flags four
`logger.exception` calls, two of them genuine misuses and two in a helper
that is only ever called from an except block.

Same change as the first half of MemMachine#1596 on speedkick, applied to main.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
malatewang pushed a commit that referenced this pull request Sep 17, 2026
…e ruff to 0.16 (#1646)

* Bring the dependency lock up to date

The lock had drifted far behind: 30 of the 52 direct dependencies were
behind their latest release, some by a major version. Refreshing it moves
112 locked packages, adds 6 and removes 2, with no downgrades. Nine are
major bumps: complexipy 6 -> 8, cryptography 49 -> 50, fastmcp 3 -> 4 (and
fastmcp-slim), mcp 1 -> 2, sentence-transformers 5 -> 6, setuptools 83 ->
84, websockets 16 -> 17, xxhash 3 -> 4. Direct dependencies that moved
include fastapi 0.139.0 -> 0.141.1, instructor 1.15.4 -> 1.17.0, langchain-aws
1.6.2 -> 1.7.6, neo4j 6.2.0 -> 6.3.1, openai 2.45.0 -> 2.54.0, qdrant-client
1.18.0 -> 1.19.0, sqlalchemy 2.0.51 -> 2.0.53, ty 0.0.59 -> 0.0.81 and
uvicorn 0.51.0 -> 0.53.0.

Two specifiers had to change for the refresh to be honest:

- litellm was capped at <1.85 when the provider was added (#1386, which
  records no reason for the bound), and dependabot later widened it to
  <1.86. That range ends inside a stretch of releases -- 1.83.8 through
  1.92.x -- that declare Requires-Python <3.14. uv ignores that upper
  bound and installed 1.85.7 anyway, but pip does not: on Python 3.14,
  `pip install memmachine-server[litellm]` resolves to 1.83.7, not the
  1.85.7 in the lock. 1.93.0 is the first release that supports 3.14,
  so that is the floor now, and the lock moves to 1.101.0. litellm also
  requires openai<3, which is what holds the OpenAI SDK on the 2.x line;
  without the floor the resolver satisfies that by walking litellm
  backwards to 1.83.0 instead.
- boto3-stubs was pinned exactly at 1.43.13 while boto3 resolved to
  1.43.47, so the stubs no longer described the installed SDK.

One source change the new versions require, and one cleanup:

- `list(vector.flat)` is a list of numpy scalars, not `list[float]`, which
  the newer numpy typing now reports at the assignment. The replay path
  converts explicitly.
- The dev group listed complexipy twice.

Left behind deliberately: nebula5-python (see the cap and its reason),
openai 3.x (litellm requires openai<3, and instructor's jiter<0.15 pin
holds openai below 3.11 regardless; 3.0 also swaps the transport to
httpx2), and ruff, still pinned at 0.15.14 -- 0.16 formats Markdown code
blocks, which would reformat six docs, and adds LOG004, which flags four
`logger.exception` calls, two of them genuine misuses and two in a helper
that is only ever called from an except block.

Same change as the first half of #1596 on speedkick, applied to main.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Drop dependencies nothing declares them for

Four entries in the server's dependency list are not requirements of this
package:

- boto3-stubs is a type-stub distribution sitting in the runtime list.
  Nothing imports it, and `ty check --project packages/server` reports
  the same "All checks passed" with the package uninstalled, so it was
  buying neither runtime behavior nor type coverage. (Base boto3-stubs
  only types the boto3 entry points; per-service clients need the
  mypy-boto3-* extras. If stronger boto3 typing is wanted later, that
  belongs in the dev group with those extras.)
- regex belongs to memmachine-common, which declares it and imports it in
  api/spec.py. Nothing in the server package uses it.
- The PyPI "dotenv" distribution ships no modules at all -- only
  dist-info -- and exists to depend on python-dotenv.
  `importlib.metadata.packages_distributions()["dotenv"]` is
  `["python-dotenv"]`, so `from dotenv import load_dotenv` has always
  been python-dotenv. Declare that instead.
- greenlet is what SQLAlchemy's asyncio extra installs, and this package
  uses create_async_engine, so ask for `sqlalchemy[asyncio]` and let it
  bring greenlet. That also covers the platforms SQLAlchemy's own
  platform_machine marker on greenlet leaves out.

The lock loses boto3-stubs, botocore-stubs, types-s3transfer and dotenv.

reranker_manager also took runtime_checkable from typing_extensions, an
undeclared dependency it reached through pydantic. It has been in
typing since 3.8 and this package requires 3.12, so it now comes from
typing, next to the Protocol the file already imports there.
memmachine-common and memmachine-client keep their typing_extensions
dependency: they support 3.10, and Self and Unpack are 3.11.

Same change as the second half of #1596 on speedkick, applied to main.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Log without a traceback where no exception is in flight

openai_embedder's dimensionality check and amazon_bedrock_reranker's
score-count check each call logger.exception() and then raise their own
ExternalServiceAPIError. Nothing is being handled at that point:
logger.exception is error(..., exc_info=True), and with no active
exception it appends "NoneType: None" to the record. logger.error is
what both sites mean; the raise that follows carries the message.

ruff 0.16 reports both as LOG004 (`.exception()` call outside exception
handlers).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Re-raise the HTTPError the client's project-error helper is given

_handle_get_project_http_error takes the caught HTTPError as a
parameter, but its bare `raise` statements and logger.exception() calls
ignored it and used the exception the *caller's* except block had
active. That works only while the helper is called from inside that
block, and ruff 0.16's LOG004 cannot see the call boundary, so it
reports the two logging calls as .exception() outside a handler.

The helper now uses what it is handed: logger.error(..., exc_info=error)
attaches the same traceback logger.exception would have, and
`raise error` re-raises the same object. Callers observe no difference:
the exception identity, its original traceback and the 422 ValueError's
__cause__ are unchanged (the helper's own frame is appended to the
traceback, as it is for any re-raise from a callee).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Move ruff to 0.16.8

ruff was held at 0.15.14 by the dependency refresh because 0.16 brings
two things that touch committed content, and both now have their own
commits ahead of this one:

- LOG004 (`.exception()` outside an exception handler), selected via
  the LOG group. Its four findings are resolved by the two preceding
  commits, so `ruff check` is clean at this pin with no suppressions.
- The formatter now formats Python code blocks inside Markdown. The six
  docs it touches are reformatted here, by `ruff format` and nothing
  else: USAGE.md, examples/v1/README.md,
  integrations/aws_strands_agent_sdk/README.md,
  integrations/langgraph/README.md, maintainers/build-pip-packages.md,
  packages/client/README.md. Trailing commas, collapsed argument
  lists, and column-aligned comments brought to two spaces; no words
  change.

The lint workflow's ruff-action reads the pin from pyproject.toml, so
CI enforces 0.16.8 from this commit on. No .py file is reformatted by
the new version.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Explain only the litellm floor in the server's pyproject

The litellm extra's comment also said litellm's openai<3 is what keeps
the resolved OpenAI SDK on the 2.x line. That describes the universal
lock, which resolves every extra, and not a requirement of this package:
the server hands the SDK no httpx client, which is the only thing
openai 3.0 changed, and its suite passes on openai 3.3.0 with the extra
absent (1869 passed, 3 skipped, the same as on the lock). Installs
without the extra may resolve openai 3.x. The comment now states only
the floor, which is this package's own decision.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

---------

Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants