Skip to content

Language grammar deprecations #3517

Description

@LiamGoodacre

We should disallow \ as an operator. This would let us commit earlier to parsing something as a lambda abstraction and potentially improve parse errors/performance. Backslash could still be used in a larger operator name, but just not by itself. So \ isn't an operator, but \\ could be.

Activity

  1. changed the title [-]Disallow backslash as an operator[/-] [+]Language grammar deprecations[/+] on Jan 31, 2019
  2. natefaubion commented on Jan 31, 2019

    @natefaubion
    Contributor

    I'm co-opting this thread to list potential deprecations for all the dark corner-cases in the current grammar.

  3. natefaubion commented on Jan 31, 2019

    @natefaubion
    Contributor

    We should disallow @ as an operator. It conflicts with binder syntax.

  4. natefaubion commented on Jan 31, 2019

    @natefaubion
    Contributor

    We should fix the precedence of @ in single binders to be consistent with multiple binders. This is currently allowed:

    do
      p@Foo a b <- go
      ...

    Where p binds over Foo a b. However in contexts that allow multiple binders, you must use parens to disambiguate.

    go p@(Foo a b) bar baz = ...

    We should always require parens in this case. This would be consistent with Haskell.


    Edit: I was slightly confused about the current state. Currently, @ binders are always lower precedence than constructor application #3532. I've just only seen it taken advantage of in single binder contexts, otherwise it breaks (as evidenced by the ticket). We should definitely just fix this.

  5. natefaubion commented on Jan 31, 2019

    @natefaubion
    Contributor

    We should fix the idiosyncrasies with layout in case branches. Currently this is allowed:

    case a of Foo bar ->
      bar + 2

    And even terribly ambiguous things like this which don't obey any reasonable layout rules:

    case a of
        Foo a ->
      true
        Bar a -> false

    We should require that case branch bodies always be indented past the pattern, just like we do in let bindings.

    The original example (which is actually somewhat common in core libs) can be replaced with:

    a # \(Foo bar) ->
      bar + 2

    And potentially include a straightforward inliner rule for this application case.

  6. natefaubion commented on Jan 31, 2019

    @natefaubion
    Contributor

    Somewhat contentious I'm sure, but I think we should disallow any whitespace besides \n, \r\n, and ASCII space (except in raw string literals). Unicode whitespace and tabs are problematic for layout, and I don't think there's a reason to support it.

  7. natefaubion commented on Jan 31, 2019

    @natefaubion
    Contributor

    We should fix precedence of kind annotated types. Currently this is allowed and I don't even...

    type Foo a b c = a :: Type -> Type -> Type b :: Type c :: Type

    This only parses because there's no kind application.

  8. hdgarrood commented on Jan 31, 2019

    @hdgarrood
    Contributor

    I suspect the single-branch case pattern is common in core libs because you didn’t used to be able to pattern match in a let binding, ie

    let Foo bar = a
     in bar + 2
    

    didn’t used to be legal, but that’s probably what I’d recommend now.

  9. garyb commented on Jan 31, 2019

    @garyb
    Member

    @hdgarrood yeah, that's exactly it.

  10. natefaubion commented on Jan 31, 2019

    @natefaubion
    Contributor

    We should disallow nested backtick expressions without parentheses. This is currently allowed:

    a `b `c` d` e

    But this is ambiguous, since it can also be parsed as:

    (a `b` c) `d` e

    Which I think is probably what one would want anyway.


    Edit: I was mistaken, the parser already works this way.

  11. natefaubion commented on Feb 19, 2019

    @natefaubion
    Contributor

    We should disallow forall as a valid identifier. It's disallowed in types, so this means that type identifiers and value identifiers have different rules. We should have a consistent rule for all identifiers.

  12. added this to the 0.13.0 milestone on Feb 24, 2019
  13. natefaubion commented on Mar 9, 2019

    @natefaubion
    Contributor

    We should simplify string and character escapes. We currently use essentially what Haskell has (really it's whatever is provided by parsec in the default language parser implementation), but it has a lot of archaic escape codes and rules. We should only allow

    • \r
    • \n
    • \t
    • \\
    • \'
    • \"
    • \x[0-9a-f]{1,4} (since we only allow utf-16 codepoints in char literals anyway)
    • \[\r\n ]+\ (whitespace gap)

    Personally, I've never even used the whitespace gap escape since we also have raw string literals, but I can see it being useful.

  14. hdgarrood commented on Mar 9, 2019

    @hdgarrood
    Contributor

    I’m not sure about restricting \x character escapes to 4 hex characters actually: that leaves you with no comfortable way of inserting astral plane characters (ie those with code point values greater than 0xffff) other than inserting the literal character. We should continue to disallow them in Char literals, since Chars are single code units rather than code points, but we should be able to treat strings as sequences of code points where appropriate, so I think “\x12345” should be valid (and equivalent to “\u{12345}” in JS).

    I’d like to also copy JS’ thing where you can use braces to delimit the hex literal from subsequent characters for clarity.

  15. natefaubion commented on Mar 10, 2019

    @natefaubion
    Contributor

    FWIW, I think you could still use surrogate pairs. Perhaps then we should allow \x with 4 hex characters, and \u{..} syntax for arbitrary escape. The problem with not having a limit on \x is you have to have some way to decide when to stop considering something as part of the escape. Haskell has \& as a zero-width escape for this purpose.

    Note that most people don't know that \& exists 😆 (I didn't even know before looking into this and @LiamGoodacre thought it was a bug that he couldn't delimit the escape).

  16. hdgarrood commented on Mar 10, 2019

    @hdgarrood
    Contributor

    Yeah I agree that we would need to have something to indicate the end of the escape, but I don’t think we should need to change the syntax in a breaking way to achieve that. My preference would be to leave \x as it is currently (which is still useful even in the absence of \&, in the case where the escape is the last thing in the string, or the subsequent character isn’t in [0-9a-f]) but additionally add braces to allow sensible delimiting going forward.

  17. natefaubion commented on Mar 31, 2019

    @natefaubion
    Contributor

    Do we want to take this opportunity to get rid of # in kind syntax. It's somewhat problematic that kinds and types have different syntax, and I think unifying them might do some good.

  18. garyb commented on Mar 31, 2019

    @garyb
    Member

    It was left that way intentionally when the other kind symbols were removed because it's unique in being the only kind that takes parameters, but I have nothing against naming it.

  19. natefaubion commented on Mar 31, 2019

    @natefaubion
    Contributor

    It was left that way intentionally when the other kind symbols were removed because it's unique in being the only kind that takes parameters, but I have nothing against naming it.

    Yes, that's true, but I think if we eventually want polykinds, it's a good idea to get rid of it. Or at least I don't think it adds anything. It's already "special" in that the language does not let you define a kind that takes a parameter.

  20. hdgarrood commented on Mar 31, 2019

    @hdgarrood
    Contributor

    Would we need to implement kind polymorphism to be able to do that sensibly though? - if the row kind constructor # were able to exist on its own, what would its kind be? Also I’d want to deprecate it for a release series, not outright remove it with little warning.

  21. hdgarrood commented on Mar 31, 2019

    @hdgarrood
    Contributor

    Also, rows are always going to be special, even if we get kind polymorphism and the ability to parameterize user-defined kinds, because they have built in support in the type checker, right? So I think it’s probably appropriate that the syntax reflects that.

  22. hdgarrood commented on Apr 1, 2019

    @hdgarrood
    Contributor

    Wait, I’m talking nonsense, aren’t I? - kinds don’t themselves have kinds.

  23. hdgarrood commented on May 8, 2019

    @hdgarrood
    Contributor

    We have tests for making sure that \ and @ are not allowed as operators and I don't think anything else in here needs additional testing, so I think this can be closed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions