Skip to content

Expose Scores for Full Text Search  #1244

Description

@lppier

Thank you for RedisGraph.
I understand that the fulltext search in RedisGraph is powered by RediSearch.
How can I get the scores (tfidf or otherwise) of the search results?
Eg.
CALL db.idx.fulltext.queryNodes('movie', 'Book') YIELD node RETURN node.title

Would it be possible to get the following?

“The Jungle Book” - score 0.8
“The Book of Life” - score 0.5

These scores would be useful for me against a threshold where I decide whether to use the node or not for further processing in my app.

Activity

  1. lppier commented on Aug 11, 2020

    @lppier
    Author

    Hi, wondering if anyone would be working on this? It would really make my current usage of redisgraph super useful. Currently we are pulling the node properties offline and searching on it.

  2. jeffreylovitz commented on Aug 12, 2020

    @jeffreylovitz
    Contributor

    Hi @lppier,

    This would definitely be a valuable addition! We've scheduled some time next week to discuss how to implement it with the RediSearch team. I'll update here when we've started development.

  3. lppier commented on Aug 13, 2020

    @lppier
    Author

    Thanks so much!

  4. lppier commented on Sep 29, 2020

    @lppier
    Author

    Hi, is this feature still under consideration?

  5. swilly22 commented on Oct 4, 2020

    @swilly22
    Contributor

    @lppier yes it is, it will be part of the next upcoming version.

  6. lppier commented on Feb 26, 2021

    @lppier
    Author

    Hi @swilly22 , can I check why this was removed from todo? Is it not technically feasible?

  7. swilly22 commented on Feb 26, 2021

    @swilly22
    Contributor

    I hope this will make it to 2.6,
    @MeirShpilraien can we allocate time to look into it next week?

  8. lppier commented on Feb 27, 2021

    @lppier
    Author

    Thanks @swilly22 , we've actually implemented our solution on redisgraph already, having the scores will allow us to use it as a feedback to our downstream results ranking module. Would really appreciate having it.

    For clarity, the request is for something like this :
    https://neo4j.com/docs/cypher-manual/current/administration/indexes-for-full-text-search/

  9. swilly22 commented on Feb 27, 2021

    @swilly22
    Contributor

    @lppier, would you mind describing your graph usecase?

  10. lppier commented on Mar 3, 2021

    @lppier
    Author

    @swilly22 We store people, companies and have relationships between them in the redisgraph. The graph is one of the backend services for our search engine. One of the use case is when someone searches for a company, say "Dropbox" , and the people affliated are also retrieved. This query is sent in parallel to the various backend services, of which redisgraph is one of them.

    Our search engine then has a ranking model which takes in a score from the various backend search services, plus the query. This ranking model decides which results to put at the top.

    Just to add, we are using the redissearch enabled wildcards like * and % in our search as well.

  11. swilly22 commented on Mar 5, 2021

    @swilly22
    Contributor

    Hi @lppier, emitting of search scores has been merged to our master branch and should be available on our edge docker image on dockerhub shortly, please see docs

  12. lppier commented on Mar 5, 2021

    @lppier
    Author

    Wow thanks @swilly22 , so it does behave like that mentioned in the docs, with a different score for each result?

  13. swilly22 commented on Mar 6, 2021

    @swilly22
    Contributor

    @lppier TFIDF describes the way results are scored

  14. lppier commented on Mar 8, 2021

    @lppier
    Author

    Hi @swilly22 , can I check when it is in the edge image, how long will it take to make it into a release?
    Thanks.

  15. swilly22 commented on Mar 8, 2021

    @swilly22
    Contributor

    it is already in the edge docker image, not sure when it will be officially released.

  16. lppier commented on Mar 9, 2021

    @lppier
    Author

    Ok thanks, am asking as I'm using it in a production system , will wait for it to be released thanks.

  17. alronlam commented on Apr 8, 2021

    @alronlam

    Hi! Would like to clarify how the Tf-Idf computation works if you're using a fuzzy query.

    That is, query looks like this: CALL db.idx.fulltext.queryNodes(‘Node’, ‘%zeller%’) YIELD node, score RETURN node, score

    Results are something like (just showing a shortened format):

    Result 1:

    displayName: Tim Keller
    score: 10
    

    Result 2:

    displayName: Bob Zeller
    score: 10
    

    I would have expected Bob Zeller to have a higher score, since there's an exact term match. Sorry, I know there's a lot of components to the score so tracing the exact cause/computation with our data is on us, but asking in case you'd have an insight into what might have caused this?

    Also, what is the range of the scores? Is it from 0-100?

    Thanks!

  18. ashtul commented on Apr 9, 2021

    @ashtul
    Contributor

    @alronlam Thank you for the issue.
    Currently, RediSearch gives all fuzzy matching terms the same score. There is no penalty on a higher distance.
    Are you familiar with an algorithm that will decide on the score based on the distance?

    Please open an issue about the subject on the RediSearch repo.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions