Skip to content

Enable exporters for resource observability, add kubelet logs in filelog receiver #509

Description

@pree-dew

@rdimitrov As per discussion here #478, adding todos:

  1. Analyse exporters for node and pod resource which include node exporter and metric server.
  2. Add filelog receiver for kubelet logs
  3. Analyse kubernetes events for different failure mode.

Activity

  1. pree-dew commented on Sep 18, 2025

    @pree-dew
    ContributorAuthor

    @rdimitrov I have started with first point.
    For 3rd point I will add the failure modes and source details first for discussion before enabling it.

  2. rdimitrov commented on Sep 18, 2025

    @rdimitrov
    Member

    @pree-dew - I think this looks good 👍 Thank you once again! 🙏

  3. pree-dew commented on Oct 8, 2025

    @pree-dew
    ContributorAuthor

    @rdimitrov Figured out the solution for all requirements here, without using any external exporter. I tried other exporters like cadvisor, kube state , metrics but was looking where with minimum moving components we can get it done.

    • Using otel collector to fetch kubernetes events and ship it as log, this will help us to figure out any issue that we see with respect to deployments etc
    • Using kubelet stats to get node, pod level resource utilisation metrics for cpu, memory, disk and network.

    I will be raising a PR once I verify these points:

    1. What is the cardinality impact of including these metrics and logs
    2. What all failure modes can be captured without missing critical ones.
    3. How frequently we want to ship this.
    4. Any unbounded label in any of the metric
  4. pree-dew commented on Oct 9, 2025

    @pree-dew
    ContributorAuthor
  5. pree-dew commented on Oct 9, 2025

    @pree-dew
    ContributorAuthor

    Pending items:

    • Specify all failure modes that will be covered
    • Prepare dashboard for resource metrics
    • Prepare dashboard for failure modes covered by k8s events.

    @rdimitrov

  6. pree-dew commented on Oct 20, 2025

    @pree-dew
    ContributorAuthor

    @domdomegg @rdimitrov Observing on staging as per the PR (#646 (comment)), so far looks good.

    Cardinality is under control and k8s events have started coming as logs. I will observe it for 1 more day so that impact of deployments on cardinality is more clear with stats.

    Image Image
  7. pree-dew commented on Oct 23, 2025

    @pree-dew
    ContributorAuthor

    @domdomegg @rdimitrov noticed that there is 1 deployment happened in default namespace in last 2 days, cardinality number looks under control so far.

    When are we planning to do a release on production? I will make sure to keep checking till that time on regular basis.

    Do you want me to send the dashboard for review or I can setup them directly on production?

  8. pree-dew commented on Oct 29, 2025

    @pree-dew
    ContributorAuthor

    @rdimitrov @domdomegg checked that production deployment gone through, everything looks okay so far. I will push small change on top of this as part of different PR :

    1. To change dot to underscore for ease of querying
    2. To reduce number of latency buckets to control cardinality.

    kubelet events have started coming as logs and resource metrics are also flowing.

  9. rdimitrov commented on Oct 29, 2025

    @rdimitrov
    Member

    Sounds good to me 👍

  10. pree-dew commented on Nov 3, 2025

    @pree-dew
    ContributorAuthor

    @rdimitrov @domdomegg Added all metrics and dashboard for the telemetry here : https://grafana.prod.registry.modelcontextprotocol.io/d/83a0f65c-bac4-40ce-a197-e77d67431ef4/registry-deployment-critical-issues

    Let me know if anything else is required, not adding alerts as of now, want to observe the usage pattern first but if you want can add for critical events added in the dashboard.

  11. rdimitrov commented on Nov 3, 2025

    @rdimitrov
    Member

    That's awesome! 🚀 Thank you! 🙏

    I think your suggestion is good and it's sensible to go with what we have now and see what alerts might be useful while we do it so we don't risk spamming ourselves 😃

    @pree-dew - I think you completed everything that was supposed to be part of this issue, so feel free to close it you think so too. Also would you like to open a separate issue for analysing and adding the potential alerts?

  12. pree-dew commented on Nov 3, 2025

    @pree-dew
    ContributorAuthor

    @rdimitrov Yes, I think separate issue for alerts and these two points would be good, I will create new issue.

    To change dot to underscore for ease of querying
    To reduce number of latency buckets to control cardinality.

    Thank you! 🙏

  13. pree-dew commented on Nov 3, 2025

    @pree-dew
    ContributorAuthor

    Created new issue here #746 , closing this one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions