Repository navigation
Language grammar deprecations #3517
Description
Activity
- changed the title
[-]Disallow backslash as an operator[/-][+]Language grammar deprecations[/+]on Jan 31, 2019 I'm co-opting this thread to list potential deprecations for all the dark corner-cases in the current grammar.
We should disallow
@as an operator. It conflicts with binder syntax.Reacted by Gary Burgess, Thomas Honeyman, Harry Garrood, Ian Jeffries, Arthur Xavier, Vasiliy Yorkin and toastalWe should fix the precedence of
@in single binders to be consistent with multiple binders. This is currently allowed:do p@Foo a b <- go ...
Where
pbinds overFoo a b. However in contexts that allow multiple binders, you must use parens to disambiguate.go p@(Foo a b) bar baz = ...We should always require parens in this case. This would be consistent with Haskell.
Edit: I was slightly confused about the current state. Currently,
@binders are always lower precedence than constructor application #3532. I've just only seen it taken advantage of in single binder contexts, otherwise it breaks (as evidenced by the ticket). We should definitely just fix this.Reacted by Gary Burgess, Harry Garrood and Rickard AnderssonWe should fix the idiosyncrasies with layout in case branches. Currently this is allowed:
case a of Foo bar -> bar + 2
And even terribly ambiguous things like this which don't obey any reasonable layout rules:
case a of Foo a -> true Bar a -> false
We should require that case branch bodies always be indented past the pattern, just like we do in let bindings.
The original example (which is actually somewhat common in core libs) can be replaced with:
a # \(Foo bar) -> bar + 2
And potentially include a straightforward inliner rule for this application case.
Reacted by Gary Burgess, Harry Garrood, Rickard Andersson and Fyodor SoikinSomewhat contentious I'm sure, but I think we should disallow any whitespace besides
\n,\r\n, and ASCII space (except in raw string literals). Unicode whitespace and tabs are problematic for layout, and I don't think there's a reason to support it.Reacted by Harry Garrood, Gary Burgess, Stefan Fehrenbach, Ian Jeffries, Matt Parsons, Rickard Andersson and Fyodor SoikinWe should fix precedence of kind annotated types. Currently this is allowed and I don't even...
type Foo a b c = a :: Type -> Type -> Type b :: Type c :: Type
This only parses because there's no kind application.
Reacted by Harry Garrood, Matt Parsons and Rickard AnderssonI suspect the single-branch case pattern is common in core libs because you didn’t used to be able to pattern match in a let binding, ie
let Foo bar = a in bar + 2didn’t used to be legal, but that’s probably what I’d recommend now.
Reacted by Hardy Jones and Gary Burgess@hdgarrood yeah, that's exactly it.
We should disallow nested backtick expressions without parentheses. This is currently allowed:
a `b `c` d` eBut this is ambiguous, since it can also be parsed as:
(a `b` c) `d` eWhich I think is probably what one would want anyway.
Edit: I was mistaken, the parser already works this way.
Reacted by Gary BurgessWe should disallow
forallas a valid identifier. It's disallowed in types, so this means that type identifiers and value identifiers have different rules. We should have a consistent rule for all identifiers.Reacted by Rickard AnderssonWe should simplify string and character escapes. We currently use essentially what Haskell has (really it's whatever is provided by parsec in the default language parser implementation), but it has a lot of archaic escape codes and rules. We should only allow
\r\n\t\\\'\"\x[0-9a-f]{1,4}(since we only allow utf-16 codepoints in char literals anyway)\[\r\n ]+\(whitespace gap)
Personally, I've never even used the whitespace gap escape since we also have raw string literals, but I can see it being useful.
I’m not sure about restricting \x character escapes to 4 hex characters actually: that leaves you with no comfortable way of inserting astral plane characters (ie those with code point values greater than 0xffff) other than inserting the literal character. We should continue to disallow them in Char literals, since Chars are single code units rather than code points, but we should be able to treat strings as sequences of code points where appropriate, so I think “\x12345” should be valid (and equivalent to “\u{12345}” in JS).
I’d like to also copy JS’ thing where you can use braces to delimit the hex literal from subsequent characters for clarity.
Reacted by Michael FicarraFWIW, I think you could still use surrogate pairs. Perhaps then we should allow
\xwith 4 hex characters, and\u{..}syntax for arbitrary escape. The problem with not having a limit on\xis you have to have some way to decide when to stop considering something as part of the escape. Haskell has\&as a zero-width escape for this purpose.Note that most people don't know that
\&exists 😆 (I didn't even know before looking into this and @LiamGoodacre thought it was a bug that he couldn't delimit the escape).Yeah I agree that we would need to have something to indicate the end of the escape, but I don’t think we should need to change the syntax in a breaking way to achieve that. My preference would be to leave \x as it is currently (which is still useful even in the absence of \&, in the case where the escape is the last thing in the string, or the subsequent character isn’t in [0-9a-f]) but additionally add braces to allow sensible delimiting going forward.
Do we want to take this opportunity to get rid of
#in kind syntax. It's somewhat problematic that kinds and types have different syntax, and I think unifying them might do some good.It was left that way intentionally when the other kind symbols were removed because it's unique in being the only kind that takes parameters, but I have nothing against naming it.
It was left that way intentionally when the other kind symbols were removed because it's unique in being the only kind that takes parameters, but I have nothing against naming it.
Yes, that's true, but I think if we eventually want polykinds, it's a good idea to get rid of it. Or at least I don't think it adds anything. It's already "special" in that the language does not let you define a kind that takes a parameter.
Would we need to implement kind polymorphism to be able to do that sensibly though? - if the row kind constructor
#were able to exist on its own, what would its kind be? Also I’d want to deprecate it for a release series, not outright remove it with little warning.Also, rows are always going to be special, even if we get kind polymorphism and the ability to parameterize user-defined kinds, because they have built in support in the type checker, right? So I think it’s probably appropriate that the syntax reflects that.
Wait, I’m talking nonsense, aren’t I? - kinds don’t themselves have kinds.
We have tests for making sure that
\and@are not allowed as operators and I don't think anything else in here needs additional testing, so I think this can be closed.
We should disallow
\as an operator. This would let us commit earlier to parsing something as a lambda abstraction and potentially improve parse errors/performance. Backslash could still be used in a larger operator name, but just not by itself. So\isn't an operator, but\\could be.