This section defines lexical details of the
TBA
language.
Reserved Sequences (Keywords)
The following character sequences are considered
reserved, and should be yield a distinct
token (rather than be counted as an identifier).
and
bool
else
false
file
if
immutable
int
or
return
sink
thrash
true
void
while
Identifiers
Any sequence of one or more letters and/or digits, and/or underscores,
starting with a letter or underscore, should be treated as an identifier
token.
Identifiers must not be reserved sequences, but may include a proper substring that is a reserved word, e.g. while1 is an identifier but while is not.
Integer Literals
Any sequence of of one or more digits yields an integer literal token as long
as it is not part of an identifer or string.
String Literals
Any string literal (a sequence of zero or more string characters
surrounded by double quotes) should yield a string literal token.
A string character is either
- an escaped character: a backslash followed by any one of the
following characters:
- n
- t
- a double quote
- another backslash
or
- a single character other than newline or double quote or backslash.
Examples of legal string literals:
""
"&!88"
"use \n to denote a newline character"
"use \" to for a quote and \\ for a backslash"
Examples of things that are not legal string literals:
"unterminated
"also unterminated \"
"backslash followed by space: \ is not allowed"
"bad escaped character: \a AND not terminated
Symbol Operators
Any of the following character
symbol sequences constitute a distinct token:
=
:
,
+
-
==
>>
>
>=
[
<<
\\(-o-)//
{
<
<=
(
!
!=
--
++
]
}
)
;
...
/
*
Comments
-
Text starting with
# up to the end of the line
is a comment (except of course if those characters are
inside a string literal).
For example:
# this is a comment
# and so is # this
# and so is # this %$!#
The scanner should recognize and ignore comments (there is no
COMMENT token).
Whitespace
-
Spaces, tabs, and newline characters are whitespace.
Whitespace separates tokens and changes the character counter,
but should otherwise be ignored (except inside
a string literal).
Illegal Characters
-
Any character that is not whitespace and is not part of a token or
comment is illegal.
Length Limits
-
No limit may be assumed on the lengths of identifiers, string literals,
integer literals, comments, etc. other than those limits imposed by the
underlying implementation of the compiler's host language.
Which Token to Produce
For the most part, the token to produce should be self-explanatory. For
example, the
+
symbol should produce the
CROSS
token, the
-
symbol should produce the
DASH
token, etc. The set of tokens can be found in
frontend.hh
or in the switch in tokens.cpp. The LCURLY token refers to a left curly brace,
{.
the RCURLY refers to a right curly brace,
}.
The lexical structure of TBA is fairly
straightforward, and largely similar to C. There are a few notable details,
some of which are departures from
C:
- The sequence
\\(-o-)// produces the THRASH token.
- The sequence
... produces the SINK token.
- The sequence
>> and << produce INPUT and OUTPUT, respectively.
- The sequences
thrash,
and
sink,
are keywords of the language. They produce the
THRASH,
and
SINK,
tokens, respectively.
- The sequence
file produces the FILE token
- The string
&& and || are NOT in the language . Instead "logical and" is represented by the string and and "logical or" is represented by the string or.
Program Behavior
Additional details (like what the behavior and syntax of
the tokens unique to
TBA will be specified as future projects approach.