Skip to content

Node Vision SDK underperforming in simple case when compared to the GCloud Vision demo #2490

Description

@maephisto

Environment details

  • OS: MacOS
  • Node.js version: 8.1.4
  • npm version: 5.0.3
  • @google-cloud/vision version: 0.11.4

Steps to reproduce

  1. download the sample image
    sample

As it is pretty clear, you'd expect it to return the capital letter "M" in the text detection field.

  1. Fire up a text editor and run the below node snippet in a env that already has GCloud authentification setup. You will also need the Vision API enabled
'use strict';

const vision = require('@google-cloud/vision')();

const options = {
  imagePath: __dirname + '/sample.png'
}

vision.detectText(options.imagePath).then((result) => {
  console.log('this is the result', JSON.stringify(result));
});

Now the result looks like this:

[
   [

   ],
   {
      "responses":[
         {
            "faceAnnotations":[],
            "landmarkAnnotations":[],
            "logoAnnotations":[],
            "labelAnnotations":[],
            "textAnnotations":[],
            "fullTextAnnotation":null,
            "safeSearchAnnotation":null,
            "imagePropertiesAnnotation":null,
            "cropHintsAnnotation":null,
            "webDetection":null,
            "error":null
         }
      ]
   }
]

If i test the same sample image on GCloud Vision landing page demo, the result is correct: the letter M
screenshot 2017-07-25 11 01 34

https://cloud.google.com/vision/

Since this is a pretty clear image and 1 single character, I'm having trouble understanding why does the accuracy differ.
I noticed this issue also on other sample images - they work in the demo but not using the Node SDK.

Activity

  1. changed the title [-]Vision SDK underperforming as compared to the GCloud Vision demo[/-] [+]Node Vision SDK underperforming in simple case when compared to the GCloud Vision demo[/+] on Jul 25, 2017
  2. stephenplusplus commented on Jul 25, 2017

    @stephenplusplus
    Contributor

    This is interesting, indeed. Thanks for reporting.

    Here is the request and response payload:

    // request
    [
      {
        "image": {
          "source": {
            "imageUri": "https://user-images.githubusercontent.com/3056582/28569874-fe84c01c-713b-11e7-8cb7-dce5e0cb10f0.png"
          }
        },
        "features": [
          {
            "type": "TEXT_DETECTION"
          }
        ]
      }
    ]
    
    // response
    {
      "responses": [
        {
          "faceAnnotations": [],
          "landmarkAnnotations": [],
          "logoAnnotations": [],
          "labelAnnotations": [],
          "textAnnotations": [],
          "fullTextAnnotation": null,
          "safeSearchAnnotation": null,
          "imagePropertiesAnnotation": null,
          "cropHintsAnnotation": null,
          "webDetection": null,
          "error": null
        }
      ]
    }

    @lukesneeringer any ideas?

  3. maephisto commented on Jul 26, 2017

    @maephisto
    Author

    Thank you for verifying this @stephenplusplus.
    @lukesneeringer Looking forward to ideas on how to overcome the issue.

  4. lukesneeringer commented on Aug 1, 2017

    @lukesneeringer
    Contributor

    Poking to say I have seen this and want to look at it ASAP.

  5. lukesneeringer commented on Aug 4, 2017

    @lukesneeringer
    Contributor

    Hey @maephisto,
    I went to troubleshoot this just now, and realized you are on an old version. We actually changed this API pretty drastically in 0.12.0; can you try it and see if you still have this problem?

  6. maephisto commented on Aug 4, 2017

    @maephisto
    Author

    Thanks for the tip @lukesneeringer did not know that 0.12 is already available!

    I've upgraded and re-tested.
    Using textDetection has the same empty result.
    But now, using documentTextDetection successfully detects the text M\n

    Below is the tested code:

    'use strict';
    
    const vision = require('@google-cloud/vision')({});
    
    vision.documentTextDetection({ source: {filename: 'sample.png' }}).then((result) => {
      // the character "M" shows up correctly
    });
    vision.textDetection({ source: {filename: 'sample.png' }}).then((result) => {
      // Result is still empty
    });
  7. added
    priority: p2Moderately-important priority. Fix may not be included in next release.
    type: bugError or flaw in code with unintended results or allowing sub-optimal usage patterns.
    on Aug 7, 2017
  8. lukesneeringer commented on Aug 8, 2017

    @lukesneeringer
    Contributor

    @maephisto Thanks. We will look at it again!

  9. added
    priority: p1Important issue which blocks shipping the next release. Will be fixed prior to next release.
    and removed
    priority: p2Moderately-important priority. Fix may not be included in next release.
    on Aug 8, 2017
  10. maephisto commented on Aug 24, 2017

    @maephisto
    Author

    @lukesneeringer thank you! any updates yet?

  11. lukesneeringer commented on Sep 11, 2017

    @lukesneeringer
    Contributor

    Okay, I finally had some time to do some investigation on this.
    I can confirm that it is not a problem with the client library, because API explorer does the same thing:

    I am sending this issue over to the Vision team to hopefully get a resolution.

  12. lukesneeringer commented on Sep 11, 2017

    @lukesneeringer
    Contributor

    document_text_detection
    text_detection

  13. added and removed
    priority: p1Important issue which blocks shipping the next release. Will be fixed prior to next release.
    on Sep 11, 2017
  14. lukesneeringer commented on Sep 12, 2017

    @lukesneeringer
    Contributor

    Okay, I got a reply back from the backend team. However, the bad news is that they basically said, "It is machine learning; not much we can do." :-/

    I recommend using DOCUMENT_TEXT_DETECTION.

  15. m3hari commented on Sep 27, 2024

    @m3hari

    Hitting this in 2024. The demo on https://cloud.google.com/vision?hl=en#demo performs much better than the SDK.

    We are using the DOCUMENT_TEXT_DETECTION+builtin/latest model. It seems there is no much we can do 🤔

    Any help or suggestion would be appreciated

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

api: visionIssues related to the Cloud Vision API.type: bugError or flaw in code with unintended results or allowing sub-optimal usage patterns.

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions