Friday, 26 August 2011

Coursework

It's done. God help me, but it's done. That means that, in theory, 60% of my workload was just ended. Now I only have two examinations to revise for and then I am free as a bird- to sort out my accomodation, health, financial, romantic (lack thereof) etc, issues.

On the plus side, I've been thinking some more about attributes. Firstly, I've decided to eliminate the different filters- now a filter is a filter is a filter. Secondly, I've decided that no matter how hard I try, it's going to be damn difficult to implement attributes on the level of the ones that ship with the language. Volatile is just the easiest example- but there are other language optimization rules, such as "Any pure function may be replaced by it's return value", or, "The implementation may automatically parallelise code marked as threadsafe", which would be difficult, if not impossible, to genuinely replicate as user-defined rules unless you wrote literally the entire optimizer and code-generator as a compile-time metaprogram- which would be impossible, because there'd be no way to code generate the code generator. Therefore, I've decided that whilst attributes offer neat semantics and relatively easy extension points to the language, the extensions that can be emplaced there are fundamentally limited. I'll have to come up with something much better if I want people to genuinely be able to write their own Const, Pure, etc.

In addition, I've been thinking some more about what I said earlier about compiling to target the JVM or CLR. Whilst I've decided again that that would be fundamentally impossible, I'm now not so sure about the reverse. As long as the Standard libraries and native interface can be dealt with, I see no reason that it should not be possible to generate "DeadMG++" code from MSIL or JVM bytecode. After all, I view them as more strict subsets of "DeadMG++" and don't see any semantics that I can't replicate. The only problem would be reflection and/or run-time code generation. Now, reflection I can work with, especially with compile-time reflection as a working base. However, I'm not sure about the ability to JIT code on the fly. I had planned to include a quite minimal JIT in the Standard library of "DeadMG++", but the JIT capabilities offered by those platforms is well in advance of what I would have put in otherwise. In addition, I would have to consider a performance loss due to a substantially less mature GC system- although compiling to native code would of course be a significant performance gain.

In addition, there's a significant library problem. Java and .NET ship with significant libraries by default- and users of .NET will expect some of the Windows-only functionality as well. This would effectively lead to me having to implement WPF, WinForms, LINQ, and everything else I can think of. Unless I could disassemble them and convert their existing implementations. That would surely be illegal... right? Certainly cool, though, if it could be done.

The final thing to consider would be interoperability- exactly how code generated from MSIL or JVM bytecode would interact with plain old "DeadMG++" code.

Saturday, 20 August 2011

Attributes- again

I've been thinking more about attributes. Fundamentally, some of them, like const, are views on a type, and some are copies. This is easily demonstrated because, quite simply, some of them are binary compatible and some aren't. That is, if you do a const_cast, then that's OK. But, if I had a Debug attribute, then the contents of a Debug std::vector may well be very different to the contents of a non-Debug std::vector, and something like a debug_cast could never exist, because there's no rules at all about what you might choose to change in a Debug std::vector. They may, in fact, be completely different types with little in common- or nearly everything in common.

An excellent example of this would be arrays. I might decide that non-Debug arrays don't check their boundaries but Debug arrays do, infact, check their boundaries. I might decide that instead of allocating Debug arrays on the native stack, I might allocate them in a separate memory arena so that they can never overflow the stack, and when an error happens, we can always get a stack trace. Hell, I might have "device" as an attribute and allocate a "device" array on the GPU or something similar- and who knows what that might involve?

That is, some attributes are like filters- you can cast them one way or another without breaking at the binary level, like const. However, some attributes are like template specialization- wholly unrelated to the source, at least, from a binary level.

This returns us to the core of the problem. What about primitive types? Users also have the option to lock user-defined types for the same effect. It's obviously illegal for user code to modify them- and most Standard types would do this, too. But I'd still want the choice to create a Tested Standard type if I wanted to.

This leads us to the inescapable conclusion that filters which do not add functionality must be valid on locked types. Furthermore, it also leads us to the next conclusion, which is that we can generalize const further to a simple removal filter.

namespace Standard {
    namespace Attributes {
        type Const {
            Type := Standard.Attributes.NonAddingFilter;
            Remove(TypeReference T) {
                // etc
            }
        }
        type Threadsafe {
            Type := Standard.Attributes.AddingFilter;
            Add(TypeReference T) {
                // etc
             }
        }
    }
}

The rules for casting are simple- the language will always implicitly reduce the functions available to the user, and always require an explicit cast to increase them. That is, NonAddingFilters are trivial to add, explicit to remove, and AddingFilters are trivial to remove, explicit cast to add, and Replacements cannot be cast at all. This means that AttributeCast will serve as a clean replacement for const_cast across any filter attributes.

For example, consider writing a webpage. You might decide logically, that an Unvalidated string would contain user input. In this case, the difference between a string and an Unvalidated string is, effectively, nothing- only that a function which writes output to the page will only take a string. The problem is that you want to have an implicit conversion in which you validate the contents. However, this violates the contract of a NonAddingFilter, namely that it doesn't add anything. You don't want to inherit from the String class, because that would produce problems like no virtual destructor, and you really don't want the implicit conversion. This is the ideal place for a Copy attribute. Copy attributes receive a type which has one member variable- the type being copied- and all functions (including constructors) appropriately forwarded. This is achievable in "DeadMG++" and not C++ because in "DeadMG++" you can go back and remove them if you don't want them, or simply start afresh. Then, all you would have to do is add a conversion operator and you're done.

namespace Web {
    type Unvalidated {
        type := Standard.Attributes.Copy;
        Result(TypeReference T) {
            T.Conversions.Add([]() {
                return Validate(this.Internal);
            });
        }
    }
}

Whilst writing up this example, I noted that if the function name is always a string then you can't access conversion operators- since the type might not even have an accessible string. This would be akin to executing arbitrary strings at compile-time. Therefore, the only way to allow a conversion to say, a compile-time type argument, would be to simply decltype() the function. This *also* yields the issue of adding functions which take compile-time arguments. If the type has an existing compile-time component, it would be inaccessible to them- that is, the only way to get compile-time functionality out of a type is to write it as a literal. I'm not too happy about this. I think it's going to be a fundamental flaw of the N-pass system that I have designed.

Secondly, the lvalue/rvalue deal. I think that the logical option is to split rvalue references and perfect forwarding. I will have one type, "Reference", which is almost like a "base class" of references. Const will not receive special treatment. A Reference may be either an lvalue or rvalue reference.

Another topic that deserves consideration is the rules for template deduction. Consider:

template<typename T> void func(T& t);
int main() {
    const int i = 1;
    func(i);
}

In C++ this will throw an error, as T cannot be deduced. In "DeadMG++", however, I am certainly considering allowing an equivalent snippet to compile. It's my understanding that C++ doesn't allow such things for legacy reasons- it would break old code. Since I don't have any old code, I feel no compunction to prevent such things from compiling. This is especially true as in "DeadMG++" I have arbitrary attributes- there would be no way to grab them all.

This still leaves me with the problem of volatile, though. The problem with volatile is not just that it behaves like a NonAddingFilter, which is fine, but it has significant implications for the code generating process- something not expressible within the attribute system.

Finally, I've realized that I've screwed myself just a little bit here- and that is, the order of functionality in a literal could matter, something I wanted to avoid. Inherently, the capability to mutate a type whilst it's still under construction is both dangerous and necessary. Consider a trivial snippet:

type T {
    SomeFunc(T) variable;
    SomeOtherFunc(T) othervariable;
}

Which of these two gains precedence? The function "SomeFunc(T)" could arbitrary mutate T in any fashion it desires. It could even lock the type. I could "lock" the type whilst it's constructing- that is, in the body of "T", then "T" is a const TypeReference, not a TypeReference.


I know that my blog posts are getting a bit infrequent and unreliable. I've got a lot going on right now and can't really afford the distraction of dreams about a better language. In a couple weeks, it'll all be over- probably for worse rather than better, but that's life.

Tuesday, 16 August 2011

C++ metaprogramming

It sucks balls. Seriously.

You can't debug it. If the compiler picks the wrong overload, good luck asking it why. You can't break in the middle of a template instantiation, and the compiler is it's usual unhelpful self. Half the time it doesn't even specify an error, it'll just say "specialization failed", or "It might be ambiguous, but it also might have no available overloads". Great, compiler. Maybe you could be more specific?

SFINAE prevents deduction. This I never knew, but apparently it actually does. How incredibly unhelpful. You might not want to compile for a reference... so I hope you didn't mind that I made your argument explicit nao.

I also discovered a need to place six bajillion typedefs in my iterator code for no apparent reason or benefit, which had me confused no end. And if you make a typo in your template code (I attempted to use a std::disable_if, which doesn't exist) the compiler gives a syntax error. I don't think so, buddy!

Oh, and then there's the declaration/definition pit of failure.

Not being able to decltype member variables is driving me nuts, since I frequently need to refer to them in decltypes for member function returns in generic code.

Want to buy: DeadMG++.

LINQ in C++

You know, LINQ really is nothing new. It's been supported in the Standard library since there was a Standard. And some trivial proxies/wrappers can prove that. The problem is just that it was butt ugly.


template<typename T, typename F> class proxy {
    T&& t;
    F&& f;
public:
    template<typename TRef, typename FRef> proxy(TRef&& tref, FRef&& fref) {
        : t(std::forward<TRef>(tref)), f(std::forward<FRef>(fref)) {}
    template<typename ret> operator ret() {
        ret retval;
        std::for_each(t.begin(), t.end(), [&](decltype(*t.begin()) ref) {
            retval.emplace_back(f(ref));
        });
        return retval;
    }
};

template<typename T, typename F> auto Select(T&& t, F&& f) -> proxy<T, F> {
    return proxy<T, F>(t, f);
}

What I really need is one: add member interface, not free function, and two, alter the lazy evaluation to be per-value on-demand, not all at once, if possible. To do this I will implement some custom iterators. Then my domination of the galaxy shall be complete. By that, I mean I will feel at least somewhat better. And less devastatingly sick.

Sunday, 14 August 2011

Tiobe language index

http://www.tiobe.com/index.php/content/paperinfo/tpci/index.html

Interesting list. The problem is that when you click on the definitions, you see that it's measured by search engine hits. Simply have more material, or old material, or wrong material, online or a more generic language name and you're "more popular". Apparently, the first 100 pages are checked for relevance. Congratulations- out of the 20 billion results for C, I'm sure that the first 2,500 provide a relevant sample.

I've been thinking more about my university work. It's so depressing. Someone please shoot me. Did I memorize the formula for generating a rotation matrix? No, I damn well did not memorize it. I looked it up on Wikipedia or I Googled it or I used the mathematical support library that comes with my API that actually puts shit on my screen. Now, I appreciate that my time has little value and I'm some nobody student who wastes it on games and languages that will likely never take off, but it sure as hell has more value than the time of Google's servers. That's free. Literally, free. For anyone. Even me.

What a total waste of my time.

Saturday, 13 August 2011

More grammatical headaches

I've got a very annoying problem in Bison, namely that it's horrendous. The tool needs functions so badly, the DRY violations are quite extreme. How many rules do I have that are

rule2
    : rule1
    | rule2 rule1

or

rule2
   : rule1
   | rule2 ',' rule1

Then we have even more fun, in the form of expression_not_identifier. That is, I need to match separate rules if the expression is an identifier or not.

It's horrendous. Now, I'm down to a single S/R conflict and it's proving a resistant one. I did successfully eliminate "::" and managed to go back to '.' for all uses.

The problem is, fundamentally, quite simple.

function: foo(x, y) {}
variable: foo(x, y) var;

Now, I *can* convince Bison to delay it by unifying the paths and some other nastiness, but that would violate DRY like politicians violate promises. As such, I'm thinking that a little script, written in Lua- the I-can't-be-bothered language of my choice- ought to handle the DRY problems. Don't you think it's ironic that I'm writing a script in Lua, to generate the input for Bison, to generate the input for the C++ compiler, which will generate a program that can generate code from another input? Meta-meta-meta-meta-meta-programming for the wincakes.

Alternatively, I could write said program in C++.

Tuesday, 9 August 2011

String literals

C++0x introduces many features to do with string literals. Now, user-defined literals, I have yet to consider for inclusion. However, I definitely did reject the extended string literals where you can pick your own delimiter and stuff like that. Why is this?

Firstly, in "DeadMG++" you can gain the string by another route anyway, like loading it from a file, reading it from a console, or even from a GUI, at compile-time, which avoids the escaping problem of C++.

Secondly, I really, really think you're doing it wrong with things like regexes. String literals are not designed to have complex semantics expressed in them. Regular expressions could easily have a more functional API, which would be vastly clearer to read for everyone.

For example:

alphabetical := Regex.Range("a", "z") || Regex.Range("A", "Z")
alphanumeric := alphabetical || Regex.Range("0", "9");
identifier := ("_" || alphabetical) 

    && Regex.ZeroOrMore("_" || alphanumeric);

match := identifier(true)(input); 

// throws if non-matching, returns match else.
if_matched = identifier(false)(input, match);

// Doesn't throw, returns bool, or puts match into "match"

They exist Under The Hood™ as FSMs anyway- this is just clearer. Secondly, this API doesn't violate DRY- whereas if you write [a-zA-Z_][a-zA-Z0-9_]*, then you specified the same alphabetical-or-underscore twice.

Of course, the above aren't especially internationalizable, which is a definite flaw. I'd have to write in Letter, Number etc as Standard.

Another thing that I've been re-evaluating is the effectiveness of compile-time/runtime deduction. I've been indecisive about how effective this can really be, but I've decided that ultimately, I should give it a shot. Partly because I feel that most code won't need to operate at both times quite like the allocator does. I am thinking right now of a kind of "supported" induction.

// usings
type AllocatorType {
// Default is private non-static non-virtual runtime

compiletime {
    HashMap(int, VariableReference) allocators;
    Default := HeapAllocator; // "malloc", essentially
    GetAllocator(Size) {
        if (Size > 52) // arbitrary for now
            return Default;
        if (allocators.Find(Size) == allocators.End())
            allocators[Size] = PoolAllocator(Size);
        return allocators[Size];
    }
}
public { // runtime default
    Allocate(Type)() {
        VariableReference Allocator := GetAllocator(Type.Size()); 

        // calls a compile-time function
        return Allocator.Allocate(Type); 

        // calls a run-time function- and returns
    }
    Allocate(Size)() {
        VariableReference Allocator := GetAllocator(Size); 

        return Allocator.Allocate(Size);
    }
    // stuff
}
}    

Having the grammar much closer to completion is really allowing me to get a better feel for coding like this. How about D's invariants? Now, the API for editing functions, I'm not too sure of right now, but worst comes to worst, I'll just directly replace it. Right now, I think that, say, adding a local variable with no name that calls the function in constructor and destructor as the first statement ought to do the trick.

// usings
Invariant(Type, Function) {
    invariant_type := type {
        Type.Pointer() This;
        type(this) {
            This = this;
            Function(This);
        }
        ~type() {
            Function(This);
        }
    };
    ForEach(Type.MemberFunction, [&](MemberFunc) {
         this := MemberFunc.Arguments["this"];
         MemberFunc.Statements.PushFront(Statement.ConstructedVariableDefinition("", invariant_type, this));
    });
    return Type;
}


This brings me back to a slight grammatical change. Previously, type definitions and type literals were different things. Now, I've decided to "promote" type definitions to full expressions. This means that you can do something like

type MyType {
    // .... some stuff here
}.Invariant([](this) {
    assert(i != 0); // semantics of "assert" TBD
});


Now, all I might have to do is create a simple "action" function which will take expressions or functions and execute them immediately...

type MyType {
    int i = 0;
public {
    SetI(value) { i = value; }
    GetI() { return value; }
}.Actions(
    [&] {

        MyType somevar;
        assert(somevar.GetI() == 0);
        assert(somevar.GetI() == somevar.i);
        somevar.SetI(5);
        assert(somevar.GetI() == 5);
    }
});



Gosh, it's like, a unit test or something! Not that I have ever written a single unit test in my entire life. But here's to hoping that arbitrary compile-time per-type actions can be used to implement unit tests.