Changelog¶
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[2.7.0] - 2026-10-01¶
Changed¶
- BREAKING (packaging): most backend-specific dependencies moved from hard requirements to opt-in extras. A plain
pip install XCoreRuntimeno longer pulls in SQLAlchemy, database drivers, Redis, Celery, Alembic, the OTLP exporter, orpython-dotenv. Each one was audited against actual usage inxcore/—ServiceContainer's per-service providers (xcore/services/container.py) already import these lazily, only when the matchingservices.*config entry is present — and moved into a matching extra if it isn't needed forimport xcoreor the zero-config boot path: XCoreRuntime[postgres]— SQLAlchemy ([asyncio]) +psycopg2, forpostgresql://URLs inservices.databasesXCoreRuntime[sqlite]— SQLAlchemy ([asyncio]) +aiosqlite, forsqlite+aiosqlite://URLsXCoreRuntime[db]— both of the above combinedXCoreRuntime[migrations]— Alembic, only used byMigrationRunner(already imported on demand, with a clearImportErrormessage if missing)XCoreRuntime[redis]— forservices.cache/services.schedulerbackend: redisortiered(the defaultmemorybackend needs none of it)XCoreRuntime[worker]— Celery, only instantiated whenservices.xworker.enabled: true(Falseby default)XCoreRuntime[tracing]— the OTLP/HTTP exporter, only imported whenobservability.tracing.endpointis set (console export, already in core, is the default)XCoreRuntime[dotenv]—.envloading, already gracefully optional in the code (try/except ImportError)XCoreRuntime[metrics]—prometheus-client, forobservability.metrics.backend: prometheus. This closes a real gap: the package was imported by core runtime code (kernel/observability/metrics.py,xcore/__init__.py) but previously only listed under dev dependencies — a production install could never actually get it.XCoreRuntime[all]— everything above, plussdk/xcli/cpp
If you relied on a bare pip install XCoreRuntime for a working database, Redis cache/scheduler, Celery worker, migrations, OTLP export, .env loading, or Prometheus metrics, add the matching extra(s). apscheduler stays in core: SchedulerConfig.enabled defaults to True, so the scheduler runs out of the box (in-memory backend) even with zero configuration — same for opentelemetry-api/-sdk, imported unconditionally at module load by kernel/observability/tracing.py.
- Dependency versions refreshed: fastapi[standard] 0.135→0.141, pydantic 2.11→2.13, sqlalchemy 2.0→2.1 (now pinned with [asyncio] in every DB extra — it was previously relying on aiosqlite to pull in greenlet transitively, which silently broke a Postgres-only install), redis[hiredis] upper bound raised 8→9, apscheduler →3.11.3, opentelemetry-api/-sdk/-exporter-otlp-proto-http 1.27→1.45, alembic →1.20, python-dotenv →1.2, psycopg2 →2.9.13, prometheus-client 0.25→0.26 (dev dependency and new metrics extra aligned to the same constraint — poetry lock rejects mismatched ones for the same package). All backend extras are duplicated into [tool.poetry.group.dev.dependencies] so poetry install --with dev (what CI runs, without --extras) still exercises every backend in tests.
Fixed¶
import xcorecrashed on a plainpip install XCoreRuntime: the extras split above was not exercised against a minimal install (CI installs the dev group, which still carries every backend).xcore/sdk/__init__.pyimportedBaseAsyncRepository/BaseSyncRepository— and therefore SQLAlchemy — at import time, andxcore.kernel.securitypullsxcore.sdkin, soimport xcoreraisedModuleNotFoundError: No module named 'sqlalchemy'and nothing could boot. Found by building the wheel and installing it in a clean virtualenv before release. The two SQL adapters are now resolved on demand (PEP 562__getattr__):from xcore.sdk import BaseAsyncRepositorybehaves exactly as before when SQLAlchemy is installed, and raises an explicitImportError(pip install 'XCoreRuntime[db]') otherwise; they are listed inxcore.sdk.__all__only when SQLAlchemy is importable, sofrom xcore.sdk import *never fails for a dependency the user does not need. Verified in a clean venv with only the base dependencies:import xcore, zero-config boot, a trusted and a sandboxed plugin loaded and called, and repeated reloads.- New regression tests (
tests/unit/test_optional_dependencies.py) simulate that minimal install in a subprocess (sys.modules[<pkg>] = Nonefor SQLAlchemy, Redis, Celery, Alembic, the DB drivers and prometheus-client):import xcore, the explicit adapter error, and a zero-config boot.
[2.6.8] - 2026-10-01¶
Fixed¶
- A healthy sandboxed worker was killed after 10 cumulative CPU-seconds, then the plugin went
FAILED:RLIMIT_CPUcounts the process's total CPU time since it started, but the worker set it once at startup (soft 10 s / hard 15 s, from a hardcoded_SANDBOX_MAX_CPU_SEC=10). The kernel's defaultSIGXCPUaction terminates the process, and no handler exists (signalis forbidden to plugins) — so any plugin doing real work was killed after a few minutes of ordinary requests, restarted, killed again, and markedFAILEDaftermax_restarts(3). Sincesandboxedis the default execution mode (2.6.2), this hit every plugin that doesn't declareexecution_mode. The limit is now two-level: - soft limit = CPU budget per request (
SandboxConfig.max_cpu_seconds, default 10): the worker re-arms it to "CPU already consumed + budget" before every call (_arm_cpu_budget,resourceis imported before the import guards go up). A request that burns more than its budget is still killed bySIGXCPU. - hard limit = lifetime ceiling of a worker generation (
SandboxConfig.max_cpu_lifetime_seconds, default 3600, +5 s grace). A hard limit can never be raised without privilege, so it stays a kernel-enforced backstop a plugin cannot lift. Before reaching it the worker recycles itself between two requests (exit codeRECYCLE_EXIT_CODE, 75) instead of being killed mid-request;SandboxProcessManagerrestarts it without counting a crash. Verified end-to-end with the real worker process: 8 requests of 0.5 s CPU each under a 1 s per-request budget all succeed (4 s cumulative — the old behavior killed the worker after the first couple); with a 4 s lifetime ceiling the worker exits with code 75 after 3 requests, not-24(SIGXCPU). Behavior change to be aware of: the kernel-enforced CPU backstop moves from 15 CPU-seconds per worker to 3600 per worker generation. Runaway requests are still bounded by the 10 s per-request soft limit and by the manager's wall-clock IPC timeout (resources.timeout_seconds, which kills and restarts a stuck worker); both settings are configurable. - A sandboxed plugin that crashed rarely still ended up
FAILEDfor good:SandboxProcessManager._restartswas only reset instart(), so three crashes spread over weeks exhaustedmax_restarts. If the subprocess ran at leastSandboxConfig.restart_reset_after(60 s) before crashing, the restart sequence now starts over; a crash loop is still capped.
Added¶
- The event loop being frozen by synchronous plugin code is now visible and attributable. A Trusted plugin runs on the application's single event-loop thread: CPU work,
time.sleep()or a blocking call between twoawaits suspends every request,asyncio.wait_for(timeout=...)cannot interrupt it (measured: ahandle()burning 1 s of CPU withtimeout_seconds: 0.2returnedstatus: okafter 1000 ms), and under the GIL threads do not help CPU-bound work.LifecycleManager.call()now times each synchronous step ofhandle()(xcore.kernel.observability.blocking.watch_blocking, transparent to results, exceptions, cancellation and timeouts) and logsplugin blocked the event loopwith the plugin, the action and the blocked time when a step exceedsplugins.loop_block_warn_ms(default250,0disables; one warning per plugin per 10 s).status()gainsmax_loop_block_ms. - Same detection for scheduler jobs (
scheduler job blocked the event loop— a synchronous job runs entirely on the loop) and for synchronous hooks that exceed theirtimeout(sync hook timed out, its worker thread keeps running:wait_forstops waiting, not the thread). SandboxConfig.max_cpu_seconds,max_cpu_lifetime_seconds,restart_reset_after;plugins.loop_block_warn_ms.
Documentation¶
doc/plugins/trusted-plugins.md: whytimeout_secondscannot interrupt synchronous code and what the new warning means; a "Reload, Unload and Memory" section (ctx.spawn_taskvs bareasyncio.create_task, what the kernel releases,plugin instance still referenced after unload, namespaced scheduler job ids); and the known limitation that a reload only affects the server worker that handled the request — there is no cross-worker reload broadcast yet.CLAUDE.md: version synced (it still said 2.5.3) and the plugin lifecycle/reload pitfalls fixed in 2.6.4–2.6.8 recorded.
Not changed (deliberately)¶
- Synchronous scheduler jobs are not moved to
asyncio.to_threadhere: it would silently change the threading model for existing jobs (no running loop in the thread, non-thread-safe state). They are reported instead; moving them is a decision for a minor release.
[2.6.7] - 2026-10-01¶
Fixed¶
- A plugin's HTTP routes were never removed on unload or reload — the old code kept serving:
Xcore._unmount_plugin_router()didapp.routes = [...], butroutesis a read-only property in Starlette (AttributeError: property 'routes' ... has no setter, swallowed by theEventBusas anevent handler error), and even with a working setter it filtered onroute.path, an attribute FastAPI ≥ 0.14x's_IncludedRouterdoes not have. A reload therefore kept the previous router mounted (the old closures served requests and pinned the old plugin generation in memory) andunload/disableleft a disabled plugin's endpoints reachable. Mounting now goes through a single_mount_plugin_router()that records the route objectsinclude_router()added; unmounting removes exactly those (in place, invalidating FastAPI's route cache) plus any route whosepathis the plugin prefix or below it. The prefix match is now segment-exact: unmounting/plugins/shopno longer removes/plugins/shop2. Boot and reload now share the same mount logic — the reload path used to always wrap the router in a second/plugins/<name>prefix while boot only did so when the router wasn't already prefixed — and the "plugin routes mounted" log no longer reads an undefinedwrapperfor a router that is already prefixed. - Unloading one instance of an Ephemeral plugin broke its sibling instances and unregistered the plugin: every pooled instance of the same manifest shared one
sys.modulesnamespace (xcore_plugin_<name>), and every instance's unload purged it and calledregistry.unregister(<plugin>). Unloading or evicting one warm-pool instance made the others' lazy relative imports fail withModuleNotFoundError: No module named 'xcore_plugin_<name>', and removed the still-active plugin from the registry until the next instance happened to re-register it. Pooled instances (LifecycleManager(..., pooled=True), used byWarmPoolandEphemeralHandler) now get their own module namespace and leave the registry alone;EphemeralHandler.stop()unregisters the plugin once, when the plugin itself stops. - A plugin that exports a service could not be reloaded after boot:
PluginSupervisor.boot()step 5 registered the entireServiceContaineras kernel-protected services — including the services plugins had just propagated into it — so the exporting plugin's next reload failed withImpossible d'écraser le service protégé '<name>' (propriétaire actuel: kernel). Services already owned by a plugin in the registry (PluginRegistry.service_owner()) are now left with their owner.
Added¶
- Forced garbage collection after unload/reload, with a leak report: a plugin instance lives in reference cycles (class, module, context, tasks), so reference counting never frees it and Python's automatic collector only reaches it after a full collection — measured on 3.14.7: a dead cycle promoted to the old generation was still alive after 11.5 million allocations on a 3-million-object heap (0 full collections ran), and the old generation of a reloaded plugin (module, state, buffers) stayed resident all that time.
LifecycleManagernow schedules onegc.collect()shortly after a persistent plugin is unloaded or reloaded (_GC_DELAY_S, 1 s; several unloads within the window share a single collection), then checks via aweakrefthat the old instance is gone. If it is not, it logsplugin instance still referenced after unloadwith the types of its referrers — the signature of a rawasyncio.create_task, a callback registered outsidectx, or a service held elsewhere. Ephemeral pool instances are exempt (short-lived young cycles; a full collection per call would cost more than it frees). A full collection is stop-the-world: ~15 ms on a 300 000-object heap, ~850 ms on 3 million. It happens once per unload/reload burst, never on the request path. Opt out withplugins.gc_after_unload: falseinintegration.yaml(defaulttrue). PluginRegistry.service_owner(name);LifecycleManager(..., pooled=...).- Regression tests:
tests/unit/test_xcore_plugin_routes.py(mount/unmount/remount against a real FastAPI app),tests/unit/kernel/test_ephemeral_isolation.py,tests/unit/kernel/test_supervisor_exported_services.py(real supervisor boot + reload),tests/unit/kernel/test_plugin_release_watcher.py.
[2.6.6] - 2026-10-01¶
Fixed¶
- Unloading a plugin left its exported services callable in the shared container:
PluginRegistry.unregister()only cleans the registry, butpropagate_services()also writes a plugin's exported services into theServiceContainer's shared dict — where they stayed after unload, reachable by every other plugin throughctx.servicesand pinning the dead plugin (instance, module, state) in memory.LifecycleManagernow remembers what it wrote and removes exactly those entries at unload — never one that another plugin has since replaced. - A plugin's HTTP router and middlewares survived unload and reload:
plugin_router/plugin_middlewareswere never reset in_do_unload(), so an unloaded handler kept the old router (and, through its closures, the old module), and a reload whose new code no longer exposed a router kept serving the previous one._collect_middlewares()also merged the newadd_state()result into the old dict (.update), so removed middlewares lived on. Both are now cleared at unload and replaced, not merged, on load. - Cancelled background tasks were not awaited before the plugin's module was dropped:
_do_unload()calledtask.cancel()and moved on, so the tasks'finallyblocks ran aftersys.moduleshad been purged. Tasks created withctx.spawn_task()are now awaited after cancellation (bounded by_TASK_CANCEL_TIMEOUT_S, 2 s; a task that swallowsCancelledErroris logged and no longer blocks the unload). An unload triggered from inside one of those tasks no longer cancels itself. ctx.spawn_task()leaked every finished task: the tracking list only ever grew for the lifetime of the plugin. Finished tasks are now dropped as they complete, and a task that fails logsspawned task failedinstead of surfacing as an unretrieved exception at garbage-collection time.
Added¶
tests/unit/kernel/test_lifecycle_unload_cleanup.py: router/middleware reset and replacement, exported-service removal (and respect for a service another plugin took over), cancelled-task cleanup ordering, tracking-list hygiene, unload from inside a spawned task.
[2.6.5] - 2026-10-01¶
Fixed¶
- After boot, the scheduler's forced cleanup silently stopped working — and so did tenant isolation for
get_service():PluginContext.get_service()consulted thePluginRegistryfirst. Duringload_all()the registry is still empty of core services, so it fell back to the plugin's ownctx.services— where the kernel had put the plugin's_ScopedSchedulertracker proxy (and the tenant-awaredb/cachewrappers when tenancy is on). But oncesupervisor.boot()finished and ranregister_core_service(), the registry started answering with the raw core objects. Any plugin loaded or reloaded after boot that didself.get_service("scheduler").add_job(...)therefore bypassed the tracker: the job stayed in_JOB_REGISTRYand APScheduler after unload, kept firing against a dead plugin, and pinned the unloaded instance, its module and its state in memory forever (verified: instance still alive aftergc.collect()). The same bypass handed out the rawdb/cacheinstead ofTenantAwareDB/TenantAwareCache, so a plugin reloaded after boot escaped tenant isolation.get_service()now serves kernel services from the plugin's own context (newPluginRegistry.is_core_service()tells kernel services from plugin exports) and keeps resolving plugin-exported services — with theirpublic/privatescoping — through the registry. - Scheduler job ids collided across plugins:
_JOB_REGISTRYand APScheduler are global and keyed by the bare job id (fn.__name__by default), so two plugins that both register acleanupjob overwrote each other, and unloading one removed the other's job._ScopedSchedulernow namespaces ids as<plugin>:<job_id>foradd_job,cronandinterval; plugins keep using their own unprefixed ids (remove_job/pause_job/resume_jobtranslate). Visible side effect: job ids shown byscheduler.jobs()and in the Redis job store are now prefixed with the plugin name.
Added¶
PluginRegistry.is_core_service(name).tests/unit/kernel/test_plugin_gc_scheduler.py: job released and instance garbage-collectable after unload after boot, no job accumulation across reloads, no collision between plugins,intervaldecorator, tenant wrapping preserved byget_service()after boot.
[2.6.4] - 2026-10-01¶
Fixed¶
- After boot, the second plugin reload/load of the whole process failed and left the plugin stuck in
FAILED:LifecycleManager.propagate_services(is_reload=True)ranself._services.update(instance_services), butself._servicesis theServiceContainer's shared dict andinstance_servicesis the plugin'sctx.servicescopy — which holds the plugin's own garbage-collection proxy (_ScopedScheduler) in place of the real scheduler (and tenant wrappers in place ofdb/cachewhen tenancy is on). The first reload therefore replaced the kernel's realschedulerin the shared container with that plugin's proxy; the next reload/load of any plugin then received a proxy-of-a-proxy, which the one-level_realunwrapping could not match against the kernel-protected service, raisingPermissionError: Impossible d'écraser le service protégé 'scheduler'. The proxy chain also grew by one level per reload. Reproduced against a realPluginRegistry+ServiceContainer(the existing tests mocked the registry, so none of this was visible):reload(A)OK, thenreload(B),load(C)andreload(A)all failed.LifecycleManagernow snapshots the services it injected (_injected_services, taken after tenant/GC wrapping, beforeon_load) andpropagate_services()ignores any entry that is still exactly that injected object — it is received, not exported — in the registry loop, in the shared-container update and in the no-registry fallback check. Only services the plugin actually exported are written back; an attempt to replace a core service with a different object is still rejected. The proxy unwrapping fallback now strips every proxy level instead of one. - A plugin in
FAILEDcould never be unloaded or reloaded: the state machine only allowedresetfromFAILED, and nothing in the kernel ever callsreset— so_do_unload()never ran for a plugin that failed mid-reload and its jobs, subscriptions, tasks andsys.modulesentries stayed registered forever.FAILEDnow also acceptsunload(forced cleanup) andreload(retry). - A failed
load()/reload()left everything the plugin had already registered in place: scheduler jobs, event/hook subscriptions, health checks, spawned tasks andsys.modulesentries of a half-initialized plugin were never released. Both paths now run the forced cleanup (without calling the plugin's own hooks, whose state is inconsistent) before raisingLoadError.
Added¶
tests/unit/kernel/test_lifecycle_reload.py: regression suite for reload/load after boot with a real registry and container (repeated reloads, cross-plugin reload, hot-load after a reload, exported services still propagated, core-service override still rejected, failed-load cleanup, unload/reload fromFAILED).
[2.6.3] - 2026-09-30¶
Fixed¶
- Sandboxed plugin logs never reached the console or
log/app.log:xcore/kernel/sandbox/worker.pyhardcoded the subprocessLOG_LEVELtoWARNINGwhen the env var wasn't set, andSandboxProcessManager._spawn()never set it — so a sandboxed plugin'sself.logger.info(...)calls were dropped at the source regardless ofintegration.yaml'sobservability.logging.level. Separately,_watch_loop()only ever read subprocess stderr once, in the crash path — so evenWARNING/ERROR-level output from a healthy running plugin sat in the OS pipe buffer and was never drained during normal operation. Sincesandboxedis now the default execution mode (2.6.2), this affected any plugin that doesn't explicitly declareexecution_mode. Fixed by threading the real configured level through a newKernelContext.log_levelfield (set fromobservability.logging.levelat boot) →SandboxedActivator→ the subprocessenvinSandboxProcessManager._spawn(), and by replacing the one-shot crash-time stderr read with a_stderr_pump()task that drains stderr continuously for the subprocess's whole lifetime and relays each line through the main process's logger (xcore/kernel/sandbox/process_manager.py).
[2.6.2] - 2026-09-28¶
Security¶
- BREAKING: a plugin whose
plugin.yamlomitsexecution_modenow defaults tosandboxedinstead oflegacy.legacyis functionally a pure alias oftrusted(PluginLoaderregisters the sameTrustedActivator()for both) — meaning an unspecified plugin was silently getting full in-process trust (no AST scan restrictions, no filesystem guard, direct access to every service) rather than the safer, more restrictive default. Flagged inreports/sandbox_dynamic_security_analysis_2026-09-28.md/roadmap/ROADMAP_PROGRESS.mdas a fail-open default worth a deliberate decision; the decision is fail-closed. Changed inxcore/kernel/security/validation.py(ManifestValidator.load_and_validate, the actual resolution path for a loaded plugin.yaml),xcore/sdk/plugin_base.py(PluginManifest.execution_modedataclass default), andxcore/sdk/manifest_schema.json(schema default + adds the previously-missingephemeralto the documented enum). Migration: any existing plugin relying on the implicit in-process default must now declareexecution_mode: trusted(orlegacy) explicitly in itsplugin.yaml, or it will load assandboxedand may fail on blocked imports/filesystem access it previously took for granted. ExecutionMode.LEGACYitself is unchanged and still resolves toTrustedActivator()when explicitly requested — only the implicit default moved.
[2.6.1] - 2026-09-28¶
Security¶
- 3 confirmed sandbox-escape techniques let a
sandboxedplugin run arbitrary commands on the host, found via dynamic testing (real plugins executed against a realXcoreinstance, not static code review — seereports/sandbox_dynamic_security_analysis_2026-09-28.md): asyncio.create_subprocess_exec/_shell—asynciowas on neither of the two forbidden-module lists.().__class__.__bases__[0].__subclasses__()walking to an already-loadedsubprocess.Popen— defeats both the static AST scan (no literal__subclasses__/import subprocessin source) and the runtime import guard (noimportstatement is ever executed; the class is already resident in memory before the guard installs).- Dynamic
import posix— listed in the static scanner'sDEFAULT_FORBIDDENbut missing from the runtime guard's_FORBIDDEN_MODULES;posixis the moduleosis built on and exposes near-equivalent primitives (fork,execve, ...). - Fixed in
xcore/kernel/sandbox/worker.py:pwd/grp/posixadded to_FORBIDDEN_MODULES, and a new guard layer patches the dangerous objects directly wherever they're reached from —subprocess.Popen.__init__,subprocess.call/run/check_call/check_output, the fullos.fork/os.exec*/os.spawn*/os.posix_spawn*/os.system/os.popenfamily, andasyncio.create_subprocess_exec/_shell— rather than only gating imports by name. This closes the__subclasses__()bypass too, which a name-based fix alone cannot. Legitimate non-subprocessasynciousage (asyncio.sleep, etc.) is unaffected — verified. - Memory limits (
RLIMIT_DATA/RLIMIT_RSS) and the filesystem guard (allowed_paths/denied_paths, directory traversal) were verified effective by the same dynamic testing — no change needed there.
[2.6.0] - 2026-09-28¶
Added¶
- Persistent plugin enable/disable state (
PluginStateStore,xcore/registry/state_store.py): until now there was no way to disable a plugin short of deleting its folder — the manifest schema'senabledfield lived underruntime.health_check, not on the plugin itself.PluginLoader.load_all()now consults a JSON file (<plugins_dir>/../.xcore/plugins_state.json) and skips disabled plugins at boot;PluginSupervisor.enable()/disable(reason=...)toggle the state live and persist it, so a process restart honors the same active/inactive set. - Forced garbage collection on unload/disable (
PluginResourceTracker,xcore/kernel/runtime/plugin_gc.py):LifecycleManager._do_unload()used to trust only the plugin's ownon_stop/on_unloadhooks plussys.modulescleanup — no scheduler job, health check, or event/hook subscription was ever unregistered, andPluginRegistry.unregister()existed but was never called. The kernel now wraps scheduler/health/events/hooks at load time to track what a plugin registers, and forces their release on unload regardless of how well the plugin's own hooks behave. Newctx.spawn_task()for background tasks that are tracked and cancelled automatically. - HTTP routes unmounted on unload:
xcore/__init__.pyonly stripped a plugin's FastAPI routes fromapp.routeson reload — never on unload/disable, leaving a disabled plugin's endpoints reachable indefinitely. Extracted into_unmount_plugin_router(), now also subscribed toplugin.*.unloaded. - HTTP control center on
/plugins/ipc/*:GET /registry(the full plugin truth table — including disabled or never-loaded plugins, whichstatus()never exposed),POST /{name}/enable,POST /{name}/disable. - IPC call supervision:
PluginSupervisor.ipc_audit()/ipc_stats()log every call (plugin,action,caller,tenant_id, status, duration) to a bounded audit trail, mirroring the existingPermissionEngine.audit_log()pattern. Exposed viaGET /plugins/ipc/audit. - Event/hook supervision:
EventBus.recent_emissions()/.stats()andHookManager.recent_emissions()keep a record of recent emissions (event, handlers/hooks matched, errors, duration) —EventBuspreviously had no metrics at all. Exposed viaGET /plugins/ipc/events.
Fixed¶
propagate_services()broke reload/re-enable of aTrustedBaseplugin:self._servicesexposes the entirectx.servicesdict for backward compatibility (including db/cache/scheduler), andpropagate_services()tried to re-register those as the plugin's own exports. This passed on first boot (the registry doesn't protect core services until afterload_all()runs), but any later reload raisedPermissionError: Impossible d'écraser le service protégé. A collision on an object identical to the one already protected (received via injection, not exported) is now ignored; a genuinely different object (an actual override attempt) still raises.- Sandboxed subprocess wasn't recycled on IPC timeout:
SandboxProcessManager.call()only caughtIPCProcessDead, notIPCTimeoutError— a subprocess that stopped responding without actually dying stayed unusable until the next periodic_health_loopcheck caught up with it.call()now recycles the subprocess immediately on either exception, on the failing request's own path, instead of waiting for the next health-check interval. PermissionEngineaudit log had no way to skip cache-hit entries: everyallows()/check()cache hit still appended to_audit_logunconditionally — the expensive part (event emission) was already skipped on cache hits, but the log append wasn't. New optionalPermissionEngine(audit_cache_hits=False)skips it; default (True) keeps the existing behavior (complete audit trail) unchanged.- Stray temp directories from crashed test runs:
tests/conftest.py'splugins_dir/temp_dirfixtures already clean up viayield+shutil.rmtree, but that teardown never runs if a test crashes hard (e.g.SIGKILL) before reaching it.temp_dirnow uses the same distinguishingxcore_test_prefix asplugins_dir, and a new session-scoped autouse fixture sweeps anyxcore_test_*directories left behind in the system temp dir at the end of the run — scoped to that exact prefix only, never a broader temp-dir sweep.
Changed¶
- Trimmed unused core dependencies:
uvicornandpydantic-settingswere never imported anywhere inxcore— both are already pulled in transitively byfastapi[standard](confirmed against its own metadata) for anyone who needs them.richhad zero usage in the core package (it belongs toxcorecli, a separate install).pydantic[email]is now plainpydantic—EmailStr/email-validatorwere never used, andfastapi[standard]already brings inemail-validatorregardless. No behavior change;poetry.lockregenerated to match.
[2.5.3] - 2026-09-09¶
Fixed¶
PluginLoader.shutdown():_stop_onecoroutine was defined but never awaited —self._handlers.clear()was called immediately, bypassing all plugin cleanup (on_unloadhooks, resource release). Now usesasyncio.gatherwith timeout per handler before clearing.AutoDispatchMixin.handle(): Scanneddir(self)on every dispatch (O(N) per call). Added lazy_build_action_map()that pre-computes adict[action_name → method]for O(1) dispatch.RoutedPlugin: Method namedRouterIn()butlifecycle.pylooks forget_router()— renamed toget_router()for consistency.PermissionEnginecache: Unboundeddictgrew forever within process lifetime. Replaced withOrderedDict-based LRU cache (max 10,000 entries) with automatic eviction of oldest entries.- TenantAware wrappers (
TenantAwareCache/DB/Scheduler):__getattr__proxy hid API surface from IDEs and mypy. Added explicit method declarations formget,mset,disconnect,ping,stats(Cache),connect,disconnect,ping,status,engine(DB),start,shutdown,health_check,status(Scheduler). - Version mismatch:
__version__.pywas2.3.3,pyproject.tomlwas2.5.2, README badge wasv2.3.5. Synced all to2.5.2. - Dead code: Removed unused
__TenancyConfigdataclass inconfigurations/sections.py. _is_db_adapter(): Fragile class-name-suffix detection replaced withisinstance()checks against actual adapter classes, with fallback for missing imports.RedisCacheBackend.clear(): Calledflushdb()which deleted the entire Redis database (not just cache keys). Now usesSCAN + DELETEto only remove matching keys.TenantAwareDB._set_tenant_schema(): Used unquoted f-string inSET search_path TO {tenant}, public. Now quotes the identifier with double quotes ("{tenant}") for PostgreSQL defense-in-depth against injection.
Changed¶
- Version synced to 2.5.3 across
__version__.py,pyproject.toml, and README badge.
[2.5.1] - 2026-08-20¶
Fixed¶
[cpp]/[all]extras removed (urgent):2.5.0shippedcpp = ["xscanner>=0.1.0"], butxscanneron PyPI is an unrelated third-party package — we never owned that name.pip install XCoreRuntime[cpp]would have silently installed a stranger's package instead of failing. Dropped both extras until the real accelerator package (xcorescanner) is published;[sdk]and[xcli]are unaffected.
[2.5.0] - 2026-08-20¶
Added¶
- Optional extras (
[project.optional-dependencies]):pip install XCoreRuntime[sdk](full plugin-author SDK,xcdk) andpip install XCoreRuntime[xcli](xcorecli, thexclicommand).[cpp]/[all]were part of this release but immediately broken — see2.5.1.
[2.4.4] - 2026-08-20¶
First release actually published to PyPI as XCoreRuntime — pip install XCoreRuntime now works. 2.4.0–2.4.3 were cut while the release pipeline itself was still being fixed and never successfully reached PyPI (see Fixed below); no functional difference to document for those beyond what's already in 2.4.0.
Changed¶
- PyPI distribution renamed to
XCoreRuntime: the registered PyPI project isn'txcore—pyproject.toml'snamedidn't match, so the project-scopedPYPI_TOKENwas rejected with403 Invalid API Token. The import name is unaffected:import xcorestill works, onlypip install <name>changes. xcoresdk/xcoreCligit dependencies dropped: PyPI rejects any package whose metadata declares a direct VCS dependency (xcoresdk @ git+https://...). Removing them brokeimport xcoreitself (ModuleNotFoundError: No module named 'sdk'—xcore/kernel/security/section.pyandvalidation.pyimportedPluginDependencyfrom the externalsdkpackage unconditionally, not just as an SDK convenience). Fixed by vendoring the pre-extraction SDK source (xcore/sdk/plugin_base.py,decorators.py,routers.py,mixin/ipc.py,adapter/*.py) back locally: the kernel now depends on nothing external, andxcore.sdk's newer features (EventMixin,HookMixin,ObservabilityMixin,ScheduledMixin,cached/cron/interval/health_check,AutoMixin, Mongo/Redis repositories) are picked up automatically if thexcdkpackage happens to be installed ([sdk]extra), and simply absent otherwise — no fake no-op fallbacks.xcore/kernel/security/{section,validation}.py:PluginDependencynow imported from...sdk.plugin_base(local) instead of the externalsdkpackage.
Fixed¶
- Release pipeline couldn't actually release:
release.ymlonly triggered onpush: tags:, but the version-bump commit + tag pushed byrelease-manual.ymluse the defaultGITHUB_TOKEN— GitHub deliberately never cascades apushevent triggered byGITHUB_TOKENinto other workflow runs (anti-loop protection).v2.3.5(era) tags were pushed with no build/publish/release ever firing.release.ymlnow also acceptsworkflow_dispatchwith ataginput, andrelease-manual.ymlexplicitly callsgh workflow run release.yml -f tag=vX.Y.Zafter pushing the tag. pypa/gh-action-pypi-publishtoken wiring: the PyPI API token was passed viaenv: PYPI_TOKEN, which the action never reads (it only reads thepassword:input) — publish step silently no-op'd on auth. Fixed towith: password: ${{ secrets.PYPI_TOKEN }}..github/workflows/labeler.yml: contained the label-mapping config ("core": - changed-files: ...) instead of a workflow definition — GitHub tried to parse it as a workflow and failed on every push. Moved the mapping to.github/labeler.yml(wherepr.yml's existing🏷️ Auto Labeljob already expected it) and deleted the broken duplicate workflow file.docs.yml: missingpermissions:block meant the Netlify PR-preview comment step failed withResource not accessible by integration. Addedcontents: read/pull-requests: write.xcore/kernel/security/validation.pyisort ordering (introduced by the SDK-vendoring fix above).
New tooling¶
.github/workflows/release-manual.yml:workflow_dispatch-only release trigger — bumppyproject.toml(explicit version orpatch/minor/major/pre-release), commit, tag, push, and dispatchrelease.yml. Supportsdry_run.
[2.4.0] - 2026-08-20¶
Added¶
- Real OpenTelemetry SDK integration and distributed trace propagation (W3C
traceparent, HTTP + sandbox IPC) — completes the tracing work started in2.3.5. Seedoc/observability/observability.md. - Tiered cache backend follow-up work, and general V2 Industrialization roadmap close-out (PR #271).
[2.3.5] - 2026-08-10¶
Closes out the V2 Industrialization roadmap: the two remaining ⚠️ items (Full OpenTelemetry, Distributed Tracing) are now implemented, and Advanced Hot Cache gets a tiered backend. V2 is now maintained in patch-release mode through December while running in production — see ROADMAP_PROGRESS.md for the V3 timeline decision.
Added¶
- Real OpenTelemetry SDK integration:
Tracer/Span(xcore/kernel/observability/tracing.py) now back onto a realTracerProviderwhenobservability.tracing.backend: opentelemetry— console export (SimpleSpanProcessor, immediate) by default, OTLP/HTTP export (BatchSpanProcessor) whenendpointis set. Public API unchanged, fully backward compatible with the previous noop implementation. - Distributed trace propagation (W3C TraceContext): a single
trace_idnow survives across process boundaries.TraceContextMiddleware(new,xcore/kernel/observability/http_middleware.py) extracts the incomingtraceparentHTTP header before any span opens; the sandbox IPC channel (sandbox/ipc.py/sandbox/worker.py) injects/parsestraceparentacross the hop to a sandboxed subprocess. Newinject_trace_context()/extract_trace_context()helpers, andspan(..., context=...)to parent a span explicitly. Tracer.shutdown(): flushes and stops theTracerProvider, wired intoXcore.shutdown(). Without it, spans still sitting in theBatchSpanProcessorbuffer at process exit were silently dropped.- Tiered cache backend (
backend: tieredinservices.cache):TieredCacheBackend(xcore/services/cache/backends/tiered.py) — memory L1 in front of Redis L2, read-through with backfill, write-through. No cross-node invalidation push (bounded byttl) — documented trade-off, seedoc/services/cache.md. - New dependencies:
opentelemetry-api,opentelemetry-sdk,opentelemetry-exporter-otlp-proto-http.
Changed¶
ServerConfig.hostdefault:0.0.0.0→127.0.0.1. Deployments that need to bind all interfaces (containers, LB in front) now do so explicitly viaapp.server.hostinintegration.yamlorXCORE__APP__SERVER__HOST.
Fixed¶
- Silent exception swallowing: 3 bare
except: passblocks now log at debug level instead of discarding the error —sandbox/ipc.py(IPCChannel.close),sandbox/worker.py(plugin.on_unload),tenancy/services.py(_set_tenant_schemacleanup). sandbox/ipc.py: replacedlogging.getLogger()with the project'sget_logger().
Documentation¶
doc/observability/observability.md: documented the real tracing backend, distributed propagation (HTTP + IPC), exporter selection table, and a new gotcha —self.tracer/self.metrics/self.healthareNoneinside ephemeral/sandboxed plugins (noPluginContextinjected there); automatic supervisor-level tracing still covers those calls without any plugin code.doc/services/cache.md: documented thetieredbackend and its cross-node staleness trade-off.
[2.3.4] - 2026-08-04¶
Added¶
- Ephemeral documentation: New
doc/plugins/ephemeral-plugins.mdguide covering what Ephemeral mode is, when to use it, the warm pool lifecycle, configuration (per-plugin + global), plugin authoring, monitoring, and tuning advice. Registered inmkdocs.yml. Updatedexecution-modes.md(3-mode comparison table + Ephemeral section),plugin-anatomy.md(manifest reference), andxcore-config.md(plugins.ephemeralsection). - CI/CD Netlify:
docs.ymlnow deploys MkDocs to Netlify instead of GitHub Pages. Production deploy on pushmain/ tags / manual, deploy preview on PR. RequiresNETLIFY_AUTH_TOKEN+NETLIFY_SITE_IDsecrets. Build is now strict (mkdocs build --strict).
Fixed¶
- Doc links: Fixed 6 broken relative links (
quickstart.md,advanced/multi-tenancy.md,sdk/examples/demo-plugin.md) that made the strict MkDocs build fail.
Fixed¶
- Ephemeral per-plugin config:
EphemeralActivatornow reads theephemeral:block frommanifest.extrawhen the manifest has noephemeralattribute (the SDK'sPluginManifestdoes not parse it as a field). Per-plugin config inplugin.yamlnow works as documented, with global fallback preserved. - warm_pool.py: Replaced
logging.getLogger()withget_logger()fromxcore.kernel.observabilityto comply with project logging conventions. Converted all 11 logger calls from%sstdlib style to structured kwargs logging.
Documentation¶
- ROADMAP_PROGRESS.md: Updated V2 progress from 70% to 85% — Ephemeral mode and Warm Pool were already implemented but marked as not done. Corrected status for Hot Cache (now ⚠️). Added "Points d'Attention" section noting duplicate
kernel/middlewares/directory.
[2.3.3] - 2026-06-08¶
Added¶
- Mode Éphémère (Ephemeral Mode): Introduced a new execution mode for plugins that optimizes RAM usage on the host machine during hot reloads. This enables fully stateless plugins and reduces resource footprints.
- Plugin Warm Pool: Implemented a warm pool mechanism to accelerate plugin activation and lifecycle transitions.
Changed¶
- Event Bus Performance: Optimized the
EventBusfor single-handler dispatch, reducing overhead for simple event flows. - Hot Reloading: Optimized the hot reloading process to be more memory-efficient by leveraging ephemeral handlers.
- Runtime Supervisor: Updated the supervisor to manage ephemeral plugin instances and warm pools efficiently.
- Internalization: Updated RBAC error messages to English for better consistency.
Fixed¶
- Resource Management: Addressed potential memory leaks during repeated hot reloads by implementing strict ephemeral lifecycle management.
[2.3.2] - 2026-06-05¶
Added¶
- Python 3.12 Support: Upgraded codebase and CI pipelines to support Python 3.12.
- C++ Security Scanner: Integrated high-performance
scanner_coreC++ extension for deeper security analysis. - Event Bus Singleton: Implemented a global
EventBussingleton available at configuration time, injected directly into middleware parameters. - Enhanced CI/CD: Added comprehensive test coverage reporting and PR size validation to GitHub Actions.
- CORS Configuration: Centralized CORS configuration in
integration.yaml.
Changed¶
- Modularization: Decoupled core runtime from SDK and CLI.
xcoreCliis now an external dependency (git+https://github.com/xcore-team/xcoreCli.git).xcoresdkis now an external dependency (git+https://github.com/xcore-team/xcoreSDK.git).
- Internal Refactoring:
- Complete overhaul of the middleware pipeline for better performance and extensibility.
- Improved database container connection handling with explicit session verification.
- Documentation: Migrated documentation system to MkDocs for better maintainability and rich search capabilities.
Fixed¶
- Plugin Sandbox: Fixed a bug where environment variables were not correctly injected into the plugin context if missing from the manifest.
- Database Reliability: Resolved an issue where database connections could fail due to unverified sessions; added automatic verification before usage.
- Plugin CLI: Fixed various bugs in plugin-related CLI commands.
[2.3.1] - 2026-05-29¶
Fixed¶
- database/session: Connections were failing silently because the session was not verified before use. Added an explicit check on the session state (
is_active) before each operation, with automatic reconnection if the session is expired or closed. - database/async_sql:
pool_pre_ping=Trueraisedping() missing 1 required positional argument: 'reconnect'when using theaiomysqldriver. Pre-ping is now disabled automatically foraiomysqlandcymysql, and is compensated by a pessimistic event listener (engine_connect) andpool_recycle. - database/async_sql: Improved handling of dead connections —
OperationalErrorandDisconnectionErrorerrors during rollback are now caught and logged instead of crashing the worker. - database/_utils: The
read_timeoutandwrite_timeoutparameters are exclusive topymysql.sanitize_connect_args()now filters them out foraiomysqlwith an explicit warning, avoiding a silent connection error. - database/migrations:
MigrationRunner._is_async()did not recognize the+aiomysqland+asyncmysuffixes, forcing the synchronous path on async connections. Both drivers are now included inasync_markers. - database/container: The
DatabaseConfigconfiguration did not expose certain production parameters (pool_timeout,pool_reset_on_return,connect_args,isolation_level,execution_options). These fields are now read fromintegration.yamland passed to the adapters.
Improved¶
- CI/CD: Updated
ci.ymlworkflow — refined the coverage step, reviewed PR labels, and added thepr.ymlworkflow to validate PR titles (conventional commits) and PR sizes. - CI/CD:
security.ymlworkflow — restricted Bandit scans to existing folders (xcore/,tests/) to eliminate false positives onextensions/andplugins/. - Tests: Fixed
test_tenancy.pytest — aligned assertion with actualContextVarbehavior after reset. - Documentation: Complete overhaul of the CLI section (
doc/cli/) with detailed guides for installation, configuration, and theworker,plugin,sandbox,manager, andmigrationcommands. Added the SDK API reference (doc/sdk/api/). - Observability: Enriched
XcoreLoggerwith structural support for contextual fields; extendedMetricsCollectorwith documentedmemoryandprometheusbackends.
[2.3.0] - 2026-05-14¶
Added¶
- Multi-tenancy Native (Axe 1):
TenantMiddleware: Extractstenant_idfrom HTTP header (X-Tenant-ID) or subdomain; injectsrequest.state.tenant_id.TenantAwareCache: Wraps cache and automatically prefixes all keys with{tenant_id}:.TenantAwareDB: Wraps SQL adapters and executesSET search_path TO {tenant_id}, public(PostgreSQL) before each query.TenantAwareScheduler: Prefixes APSchedulerjob_idwith{tenant_id}:.wrap_services_for_tenant(): Replaces services in plugin context at each call; zero code changes for existing plugins.
- IPC Authorization (allowed_callers):
IPCAuthMiddleware: First middleware in the pipeline; checksallowed_callersdeclared inplugin.yaml.- Deny-by-default: IPC calls are denied if the list is empty or missing. Direct HTTP calls (caller=None) still pass.
PluginLoader.get_manifest(name): Added method to retrieve manifest from middleware.
- @schema Decorator (Axe 3):
- Versioned decorator with built-in validation (Pydantic).
SchemaRegistry: Singleton storing all schemas declared via@schema.BreakingChangeDetector: Detects breaking changes between two registry versions.- CLI:
xcore plugin validate --check-breaking schemas_v1.json.
- Configuration:
tenancy:section inintegration.yamlwith 8 flags:enabled,header,subdomain,default_tenant,isolate_cache,isolate_db,isolate_scheduler,enforce_ipc.TenancyConfigdataclass inconfigurations/sections.py.allowed_callers: list[str]added toPluginManifest.
- Testing:
- 58 new tests:
tests/unit/kernel/test_tenancy.py(41) andtests/integration/test_tenancy_integration.py(17).
- 58 new tests:
- Documentation:
doc/guides/tenancy.md: Complete multi-tenant guide.doc/guides/plugin-manifest.md:plugin.yamlreference.doc/reference/configuration.md: Documentedtenancy:section.doc/reference/sdk.md: Documented@schema.doc/guides/security.md: IPC andallowed_callerssection.doc/architecture/decisions.md: Decisions 7 (location), 8 (IPC deny-by-default), 9 (@schema source of truth).
[2.2.1] - 2026-05-24¶
Fixed¶
- database/async_sql:
pool_pre_ping=Truecausedping() missing 1 required positional argument: 'reconnect'with aiomysql. Pre-ping is now disabled automatically for aiomysql/cymysql and compensated by a pessimistic event listener +pool_recycle. - database/migrations:
MigrationRunner._is_async()did not recognize+aiomysqland+asyncmydrivers, forcing synchronous path on async connections. - database/_utils:
read_timeoutandwrite_timeoutare pymysql-only parameters.sanitize_connect_argsnow filters them for aiomysql with an explicit warning.
[2.2.0] - 2026-05-24¶
Added¶
- DatabaseConfig: New configurable pool parameters in
xcore.yaml:pool_pre_ping,pool_recycle,pool_timeout,pool_reset_on_return,connect_args,isolation_level,execution_options. - database/adapters/_utils.py: New module for driver detection and connection argument sanitization.
Fixed¶
- database/async_sql: Fixed stale connections (MySQL/MariaDB) after
wait_timeout. - database/async_sql: Added missing
@asynccontextmanageronsession(). - database/async_sql + sql: Added missing
disconnect(). - database/async_sql + sql: Improved error handling during rollback on dead connections.
[2.2.0] - 2026-05-14¶
Changed¶
- Security: Removed
python-joseandpython-ecdsato eliminate vulnerability to Minerva timing attacks (CVE-2024-23342). - Cleanup: Removed 7 unused dependencies (
pillow,watchdog,user-agents,aiocache,toml,mysql-connector-python). - Optimization: Moved
psutilto dev dependencies andmarkdownto docs dependencies.
[2.1.3] - 2026-05-13¶
Added¶
- XWorker (Native Celery): Full Celery integration in
ServiceContainer. - CLI xcore worker: Command to manage FastAPI and Celery processes (
start,stop,status,logs, etc.). - Extended Configuration: FastAPI constructor parameters and uvicorn parameters configurable via YAML.
- Declarative Middleware System: Automatic loading from
integration.yaml.
[2.1.2] - 2026-04-29¶
Fixed¶
- 13 critical test failures resolved (kernel, permissions, sandbox).
- AST Scanner: detection of bypasses via import aliases.
Improved¶
- Performance:
- LRU Cache on
PermissionEngine: +34% throughput. - Native
mset/mgeton Redis: up to 77x faster on batch operations. - Pre-compiled regex in
Policy.matches(): short-circuit in 0.4 µs.
- LRU Cache on
- Quality:
pytest-benchmarkintegration.- Pre-commit hooks for black, isort, and flake8.
pyproject.tomlmigrated to PEP 621.
[2.0.0] - 2026-04-15¶
Added¶
- Plugin-First Architecture: Modular kernel, separation of Kernel / Services / Plugins.
- Advanced Sandboxing: OS subprocess isolation, JSON-RPC 2.0 communication.
- ServiceContainer: Dependency injection for DB (SQLAlchemy 2.0), Cache (Redis/Memory), Scheduler (APScheduler).
- MiddlewarePipeline: Pre-compiled pipeline (Tracing → RateLimit → Permissions → Retry).
- SDK:
@action,@router,@validate_payload,AutoDispatchMixin,RoutedPlugin. - RBAC: Pluggable
AuthBackend+ declarativeRBACChecker. - StateMachine: FSM per plugin with validated transitions.
- PluginRegistry: Metadata, dependencies, semver versioning.
[1.x] - Legacy¶
Added¶
- Initial stable release based on FastAPI.
- Monolithic plugin system without isolation.
- Limited support for asynchronous services.