Skip to content

Consider removing napi_get_value_string_length #226

Description

@jasongin

I regret adding this API now, because it doesn't serve much purpose, and is only likely to cause confusion and bugs. The intention was that this would return the number of "characters" in a string independent of encoding, but that's not generally useful. In almost all cases, one of the encoding-specific napi_get_value_string_* APIs is more correct. (Pass a null buffer if only the encoded length is desired.)

Anyway the current V8 implementation of napi_get_value_string_length() is technically wrong: it returns the number of 2-byte code points of the UTF-16 encoding, but there are actually characters that are encoded as two UTF-16 code points.

Activity

  1. kkoopa commented on Apr 12, 2017

    @kkoopa

    You are probably right. It is better to err on the side of caution. Character encoding should be added to the list of 2 hard problems in CS, which would now be the 3 hard problems:

    1. Cache invalidation
    2. Naming
    3. Character encoding
    4. Off-by-one errors
  2. jasongin commented on Apr 12, 2017

    @jasongin
    MemberAuthor

    And in case anyone is wondering, the JavaScript String.prototype.length property returns the number of UTF-16 code units, which may be different from the number of characters. So, getting the true character count is not common with JavaScript, and is probably best left to specialized internationalization libraries.

  3. trevnorris commented on Apr 14, 2017

    @trevnorris

    @jasongin Personally I think the byte length would be the most useful b/c that's what's needed if you want to copy out the string. For that you could just copy the method used by Buffer.byteLength(), and in the same way could extend the API to accept an encoding.

  4. jasongin commented on Apr 14, 2017

    @jasongin
    MemberAuthor

    @trevnorris There are other APIs to get byte length, that is why this API should just be removed.

    napi_get_value_string_utf8() gets the length of the UTF-8 encoding of a string, in bytes.
    napi_get_value_string_utf16() gets the length of the UTF-16 encoding of a string, in 2-byte code units.
    In both of those, if you just want the length without copying the string, you can pass in a null string buffer.

  5. mhdawson commented on Apr 18, 2017

    @mhdawson
    Member

    I'm +1 for removing. Its safer to start with a smaller API and add back than to be stuck with methods that cause confusion and are not really needed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      Sponsor
      SponsoredKunjungi sekarang
      Promo