Wednesday, March 19, 2014

Speculative stores

Recently I read this wonderful article on C11 atomic variables and its possible usage in Linux Kernel. This is why admire community coding. People throng into fruitful discussions and eventually best comes out. The final result is lot of learning from all corners. In the article, Mr.Corbet mentioned about speculative stores which means the consequences of nasty compiler optimizations. I wrote about how intelligibly compilers handle certain cases here and here. So lets look at following example which is also mentioned in the article.

int y=2;

int do_some_work()
{
    y = 2;

    if (y)
        ....
    else
        ....
}


In above code, many programmers may expect compiler to rip off the 'else' branch. That's the dangerous part if the code belongs to kernel space. Why? Now 'y' is a data segment entity and can be manipulated by any CPU in SMP system. It can be even set to zero by some CPU. In that case, the optimization of compiler will result in untidy results. I cannot say how do compilers treat such code in kernel space since it requires bit of time to experiment.

How can we do it in user space? Very simple! Run more than one thread and lets see how assembly looks like. Note that this code is not thread safe

#include<pthread.h>
#include<stdio.h>
#include<assert.h>

int y=0;
void* thread_routine(void* arg);

void* thread_routine(void* arg)
{
        y=1;
        if(y)
                printf("Y is Y in thread = %d\n", pthread_self());
        else
                printf("Y is !Y in thread = %d\n", pthread_self());
}

void* thread_routine2(void* arg)
{
        y=0;
        if(y)
                printf("Y is !Y in thread = %d\n", pthread_self());
        else
                printf("Y is Y in thread = %d\n", pthread_self());
}

int main(int argc, char **argv)
{
        pthread_t tid[2];

        int thread_rc = 0;

        thread_rc = pthread_create(&tid[0], NULL, thread_routine, NULL);
        assert(!thread_rc);
        thread_rc = pthread_create(&tid[1], NULL, thread_routine2, NULL);
        assert(!thread_rc);
}


I am stripping of unnecessary sections and retaining only the thread stack assembly. The thread_routine2 function also looks similar

thread_routine:
.LFB2:
        .cfi_startproc
        pushq   %rbp
        .cfi_def_cfa_offset 16
        .cfi_offset 6, -16
        movq    %rsp, %rbp
        .cfi_def_cfa_register 6
        subq    $16, %rsp
        movq    %rdi, -8(%rbp)
        movl    $1, y(%rip)
       movl    y(%rip), %eax <-- I know you are tricking me :D
       testl   %eax, %eax
       je      .L2

        call    pthread_self
        movq    %rax, %rsi
        movl    $.LC0, %edi
        movl    $0, %eax
        call    printf
        jmp     .L4
.L2:
        call    pthread_self
        movq    %rax, %rsi
        movl    $.LC1, %edi
        movl    $0, %eax
        call    printf

.L4:
        leave
        .cfi_def_cfa 7, 8
        ret
        .cfi_endproc


If you glance at assembly code, gcc is smart man :D. It knows that it should not optimize in such cases. If you observe the assembly, gcc emits code for both if and else part even though there is straight forward assignment before the branching. Also look at "movl y(%rip), %eax"! Instead of blindly copying value of '1' to EAX register, the actual value of 'y' is copied and tested :-).

Caveat: Multi threading may not emulate a SMP scenario in linux. Nowadays operating systems tend to hook threads to particular CPU rather than multiple CPUs. This is mainly to avoid penalty incurred due to cache line invalidations especially when global variable is involved and can be modified. Nevertheless, a thread can be pre-empted in middle of operation (say while if{} branch can be precisely evaluated) unless lock is held explicitly. Understanding SMP systems is quite intricate however opens up to wide variety of thoughts in programming world. Two cores are not two brains you know ;-). There are lot difficulties while handling such scenarios!

Finally short assignments ;-): 

1) Examine the assembly in case of -O2 switch
2) Remove threads and run bare minimal program while retaining data segment  variable and observe what gcc does!

Thursday, March 6, 2014

Handling of absurd code by gcc - The absurd sequel ;-)

I did mention about this in my previous blog here. Now slightly extending the condition to this

int test(unsigned int k)
{
        if (k <= 0)
                printf("BUG IN COMPILER: THIS IS ABSURD BLOCK\n");
}


We can expect few changes by compiler. Now this has two conditions, one for comparing for zero and other for comparing for Sign bit. What does compiler emit?

.LFB0:
        .cfi_startproc
        pushq   %rbp
        .cfi_def_cfa_offset 16
        .cfi_offset 6, -16
        movq    %rsp, %rbp
        .cfi_def_cfa_register 6
        subq    $16, %rsp
        movl    %edi, -4(%rbp)
        cmpl    $0, -4(%rbp)
<-- Just compare with zero and kick programmer 
        jne     .L3
        movl    $.LC0, %edi
        call    puts



As expected it emits only condition for comparing with zero :D. gcc is very smart with this case too! Everything remains same except we have "call puts" now to print in case of condition is true.

Wednesday, March 5, 2014

Handling of absurd code by gcc

We programmers are bound to make silly mistakes and tend to write absurd code :-). If these things are spotted during code reviews or internal testing, then you are spared. If the buggy stuffs land in customer's runway, then we are in trouble :-). Today's blog is to closely examine gcc behavior's to such absurd code. This example is not exhaustive. I will try to come up with few more examples in future to learn myself and share observations. As of now, here is sample code.

#include<stdio.h>

int test(unsigned int k)
{
        if (k < 0)
                printf("BUG IN COMPILER: THIS IS ABSURD BLOCK\n");
}

int main()
{
        unsigned int k = 9;
        test(k);
        return 0;
}


As a programmer you know the absurdity of the code. So how does gcc behave. We can only say by looking into assembly output emitted by gcc. Here is the assembly dump of the program (with default optimization).

<snip>

test:
.LFB0:
        .cfi_startproc
        pushq   %rbp /* Save return pointer */
        .cfi_def_cfa_offset 16
        .cfi_offset 6, -16
        movq    %rsp, %rbp /* New base pointer */
        .cfi_def_cfa_register 6
       movl    %edi, -4(%rbp) /* Copy Argument */
       popq    %rbp /* Restore base pointer */
        .cfi_def_cfa 7, 8
        ret
        .cfi_endproc
.LFE0:
        .size   test, .-test
        .globl  main
        .type   main, @function
main:
.LFB1:
        .cfi_startproc
        pushq   %rbp
        .cfi_def_cfa_offset 16
        .cfi_offset 6, -16
        movq    %rsp, %rbp
        .cfi_def_cfa_register 6
        subq    $16, %rsp
        movl    $9, -4(%rbp)
        movl    -4(%rbp), %eax
        movl    %eax, %edi
       call    test /* Here is call to function */
        movl    $0, %eax
        leave
        .cfi_def_cfa 7, 8
        ret


<snip>


As you can see, the entire 'if' block is discarded by gcc :-D. Compilers are smart nowadays ;-). They know how to get rid of weeds and make ELF fertile :-). Even though we inject irrelevant code, compilers (atleast gcc) get rid of them in final binary. Let me see what more weird stuff can be experimented with. Hope you enjoyed this small and simple post. Critiques and comments are always welcome!

Saturday, February 22, 2014

LD_PRELOAD environment variable - A short insight

I believe most of the programmers (unix C programmers) are aware of the application of LD_PRELOAD env variable. Basically it can be used to override any functionality of libc with the in house implementation. For ex: 'getaddrinfo' may have different implementation for an organization and cannot replaced in libc code itself (due to GPL restrictions). The organization may not be wanting to expose their internal implementation to public domain to preserve intellectual stuffs. Can't they write their own implementation? That is crux of the matter. They may be using third party application and want to modify certain functionalities for their use. Instead of modifying glibc code, they replace with their own implementation. Aha! how is this irony :-). Forget it! The usual method is to implement same implementation with prototype matching libc and create a dynamic shared library. Later set the LD_PRELOAD environment variable to this dynamic shared object. This makes the program loader to preload the created .so before libc gets loaded. The loader marks all the symbols used by executable based on the order of libs loaded. In above case, 'getaddrinfo' will be linked to created library than libc. Even though I knew this information, never attempted to write a program. Finally I was tempted to write and here it is. This guy overrides malloc implementation

The library for malloc -- Buggy!!!

my_lib.c

#include <stdio.h>

char *virus_data_segment = (char*)(0x1234567890)

void* malloc(size_t size)
{
        printf("Get lost! I do not have even a penny to give ;-)\n");
        return virus_data_segment;
}


created shared library: gcc -shared -o libmy_lib.so -fPIC my_lib.c

Now the actual fun starts :D

set LD_PRELOAD

>>LD_PRELOAD=./libmy_lib.so

execute ls command

>> ls
Get lost! I do not have even a penny to give ;-)
Segmentation fault


LOL! The malloc has been overridden and 'ls' cannot run (may be 'ls' is using malloc). Why segmentation fault? Because malloc returned an address region not belonging to process address space. Here is output when we return NULL.

Get lost! I do not have even a penny to give ;-)
Get lost! I do not have even a penny to give ;-)
Get lost! I do not have even a penny to give ;-)
Get lost! I do not have even a penny to give ;-)
Get lost! I do not have even a penny to give ;-)
Get lost! I do not have even a penny to give ;-)
Get lost! I do not have even a penny to give ;-)
Get lost! I do not have even a penny to give ;-)
Get lost! I do not have even a penny to give ;-)
ls: memory exhausted


Strange is it not! Either it is retrying or multiple places trying to allocate.

Yes, LD_PRELOAD is well known :-) but scribbled here in case for anyone can provide more insight. I will try to gather some more information if possible and post them.

Thursday, December 12, 2013

The 'cd' command

As everyone knows, the 'cd' command in Linux changes the current working directory to new directory. Really?! To be precise it changes the current working directory of the process in interest to new working directory. It cannot change 'cwd' of other processes. So whats the big deal?

Internally 'cd' command uses chdir system call to change current working directory. As mentioned before, 'chdir' can only change 'cwd' of process which is calling it not any other process. So what? :-). Now think of bash which is executing 'cd' command. How does it change its 'cwd' to new directory? Usually other commands like 'ls', 'dir' etc.. are executed with fork+exec combination i.e. by spawning a new process. Now in case of 'cd' you cannot spawn new process since new process cannot change 'cwd' of bash.

Aha! there is interesting part. Now how do you work it around? What does bash do? Bash does this by embedding the implementation of 'cd' in its own executable i.e. 'cd' is a command in bash itself rather than being stand-alone executables like 'ls' or 'dir'. Bash implements 'cd' in itself and exposes it as command in terminal. The user still interprets it as stand alone command because of this bash trick ;-). That means along with other commands enumerated by bash (using PATH variable), it also inserts 'cd' into the pool. Since 'cd' is now part of bash process, the changing to new working directory is straight forward :-). You can check your bin directory if any executable with name 'cd' could be found like one below ;-)

nandakumar@heramba ~ $ which ls || echo -e "get lost :-)"
/bin/ls
nandakumar@heramba ~ $ which cd || echo -e "get lost :-)"
get lost :-)

Even I was not aware of this fact until I recently read System Programming Book by Robert Love. The beauty of book is how Mr.Robert Love presents such minute things so accurately. At the end of day, there was a happy learner!

Thursday, December 5, 2013

Thread local storage with gcc __thread keyword

gcc provides __thread keyword to make a global variable (or in general data segment variable) local to thread.

This may be required when you use want to use thread safe/specific data within your code.

Consider an example:

#include <stdio.h>
#include <pthread.h>

void iterate()
{
        static int i=0;
        i++;
        printf("Thread id: %x, i=%d\n", pthread_self(), i);
}

void* thread_func (void* data)
{
        iterate();
}

int main()
{
        pthread_t tid[5];
        int i=0;

        for (i=0; i<5; i++)
                pthread_create(&tid[i], NULL, thread_func, NULL);

        for (i=0; i<5; i++)
                pthread_join(tid[i], NULL);
}


Here is output:

Thread id: 6ebcc700, i=1
Thread id: 6e3cb700, i=2
Thread id: 6d3c9700, i=4
Thread id: 6cbc8700, i=5
Thread id: 6dbca700, i=3


Static will be part of data segment which is not thread safe. We can make it thread safe by adding __thread keyword. Here is modified snippet of program and rest all things remain same.

<snip>

void iterate()
{
        static __thread int i=0;
        i++;
        printf("Thread id: %x, i=%d\n", pthread_self(), i);
}

<snip>


And the output:

Thread id: 31339700, i=1
Thread id: 2f335700, i=1
Thread id: 30b38700, i=1
Thread id: 30337700, i=1
Thread id: 2fb36700, i=1

As far as I know, __thread keyword can only be used with POC types (Plain old C types) but not hybrid or pointer types (citation needed). In that case, next statement provides an answer to achieve it. Also there is obvious overhead using __thread keyword since it requires some internal manipulation to get the data of particular thread of interest. (A simple dig into the assembly code will reveal IMHO)

There are also pthread_getspecific() and pthread_setspecific() APIs POSIX provides for TLS. Will try to experiment on the same in future.

Wednesday, July 10, 2013

Typechecking with gcc typeof

Some more fun with gcc! More and more I look into linux kernel source, more I explore the new things about gcc. As far as I know, most of gcc specific implementations are driven by linux kernel community which is replicated by other proprietary compiler writers and finally becomes ISO C standard :-). That is power of open source!

Here is a macro defined in one of linux kernel source header. We know that strict type checking is standard by itself in C++. However, since most of C clients are driver writers, it is still maintained as lightweight. But you can approximate the typecheck with following macro if required. This generates warning by default if the types do not match and can be transformed to error with -Werror switch too! One more new gcc specific implementation I learnt was the compound statement within parenthesis.

#define typecheck(type,x) \
({  type __dummy; \
    typeof(x) __dummy2; \
    (void)(&__dummy == &__dummy2); \
    1; \
})


Let me dig into the macro.

The first line is normal declaration.
Second line gets datatype of 'x'
Third line compares the address of the pointer types.
Fourth line is the final value of compound statement.

There are intentions behind lines #3 and #4. The line 3 is just for generating warning of invalid comparison of types and not meant for any logical matching. The 'void' signifies to ignore the comparison value. The final '1' signifies always success. The value evaluated from compound statement is the final value evaluated in compound statement. If none is present, the evaluation is treated as void. The above macro is just for comparing types. This is reason why the evaluation of pointer comparison is masked off and final expression is *always* made to return true value.

Here is slightly different macro and associated example. I have used it to typecheck pointer types. This may be handy while checking for void* arguments in functions.

#define typecheck(type,x) \
({  type *__dummy=NULL; \
    typeof(x) __dummy2; \
    (void)(__dummy == __dummy2); \
    1; \
})

struct test_struct {
        int k;
};

int main()
{
        int y = 10;
        typecheck(struct test_struct*, &y);

}


Here is sample output from gcc.

nandakumar@HERAMBA ~/codes $ gcc test.c
test.c: In function ‘main’:
test.c:15:2: warning: comparison of distinct pointer types lacks a cast [enabled by default]


If you want permanent error here is output from gcc with -Werror switch

nandakumar@HERAMBA ~/codes $ gcc -Werror test.c
test.c: In function ‘main’:
test.c:15:2: error: comparison of distinct pointer types lacks a cast [-Werror]
cc1: all warnings being treated as errors


This macro is really handy and can be replaced with compiler specific implementations as a one shot change. Note that macro implementation may not be same across all compilers.

Extra Notes:

1) The typeof operator and compound statement within parenthesis seems to be gcc and GNU C specific.

2) There is always a great thing about the compound statement within parenthesis. The variable names used inside the statement may also clash with the names being used within main code block which is not possible with plain macros. For example, the below code still works fine!

#define typecheck(type,x) \
({  type *__dummy=NULL; \
    typeof(x) __dummy2; \
    (void)(__dummy == __dummy2); \
    1; \
})

struct test_struct {
        int k;
};

int main()
{
        int __dummy = 10;
        typecheck(struct test_struct*, &__dummy);

}

And here is output!

nandakumar@HERAMBA ~/codes $ gcc test.c
test.c: In function ‘main’:
test.c:15:2: warning: comparison of distinct pointer types lacks a cast [enabled by default]


References:

1) typeof: http://gcc.gnu.org/onlinedocs/gcc/Typeof.html
2) Compound statements: http://gcc.gnu.org/onlinedocs/gcc/Statement-Exprs.html#Statement-Exprs