Showing posts with label Articles. Show all posts
Showing posts with label Articles. Show all posts

Wednesday, April 15, 2009

Open CV

OpenCV means Open Source Computer Vision Library. It is a collection of C functions and a few C++ classes that implement many popular Image Processing and Computer Vision algorithms.

The key features:

OpenCV provides cross-platform middle-to-high level API that includes about 300 C functions and a few C++ classes. Also there are constantly improving Python bindings to OpenCV, see interfaces/swig/python and samples/python. OpenCV has no strict dependencies on external libraries, though it can use some (such as libjpeg, ffmpeg, GTK+ etc.) when it is possible.


it can be downloaded from:

http://www.sourceforge.net/projects/opencvlibrary

Wednesday, March 18, 2009

The art of Decompilation

Decompilation


Decompilation is the reverse process of compilation i.e. creating high level language code from machine/assembly language code. At the basic level, it just requires to understand the machine/assembly code and rewrite it into a high level language, but things are not as simple as they seem, particularly when it comes to implementing a decompiler. Throughout this discussion, we will be using the C language for the high level language, and the 8086 assembly language for the low level language.

The ethics of decompilation

Is decompilation legal, and is it allowed?

There are many situations when decompilation can be used...

  1. To recover lost source code. You may have written a program for which you only have the executable now (or you got the exe of a program you wrote long back, from someone else!). If you want to have the source for such a program, you can use decompilation to recover it. In all rights, you are the owner of the program, so nobody is going to question you.
  2. Just as stated above, applications written long back for a legacy computer may not have the source code now, and you may need to port it to a new platform. Either you have to rewrite the application from the scratch, or use decompilation to understand the working of the application and write it again.
  3. Say you have code written in some language for which you cant find a compiler today! If you have the executable, just decompile it and rewrite the logic in the language of your choice today.
  4. To discover the internals of someone else's program (like what algorithm they have used...)

Usually all software are copyrighted by the authors. This means, copying or expressing the same idea in another program is prohibited. Hence if you are using decompilation to discover the internals of a program and if that particular part is breaching the copyright of the owner, you are liable for legal action. However, there are some permitted uses of decompilation, like the first three cases stated above. Also, decompilation of parts of software which do not come under the copyright laws (e.g. algorithms) is permitted. In any case, it is better to contact your legal advisor if you are doing any serious work with decompilation.

In all practical purposes, decompiling programs which were created by you can't be questioned! After all, you are the owner of all rights to the program. But be careful if you are trying it out on someone else's programs.

Is decompilation possible?

Let's take a look at a normal C compiler. When a C program is compiled, the first stage of the compiler will generate a very rudimentary assembly language output (or nearly equivalent to it), which is nothing but a line-by-line translation of the C source code. If any optimizations are chosen to be done, the next stages perform the code optimization to replace redundant instructions and improve the overall efficiency of the output program. This output is then linked with the standard libraries for any library function calls, and saved in the executable format of the platform.

void main()
{
int i, j, k;

i = 10;
j = 20;

k = i*j + 5;
}

 ;---- i = 10;
mov [bp+2], 10
;---- j = 20;
mov [bp+4], 20
;---- k = i*j + 5;
mov ax, [bp+2]
mov bx, [bp+4]
mul bx
add ax, 5
mov [bp+6], ax
Sample C code
Possible output of a compiler
(without optimizations)

If no optimizations are performed while the code is generated, it is very easy to understand what the output code does, and an equivalent C code can be written/generated automatically. Note that we can only generate "an equivalent C code", and not the "same C code" which was compiled to get this executable. In other words, it is always impossible to get the exact source code, but we can generate an equivalent program which will function in the same way.

However, things are a bit different if any compiler optimizations are used while building the original executable. It turns out to be more difficult to understand the flow of the program now, and the more rigorous the optimizations are, the worse are our chances to figure out what the code is doing exactly.

void main()
{
int i, j, k;

i = 10;
j = 20;

k = i*j + 5;
}

 ;---- i = 10;
mov si, 10
;---- j = 20;
mov di, 20
;---- k = i*j + 5;
mov ax, si
mov bx, di
mul bx
add ax, 5
mov [bp+6], ax

 ;---- k = i*j + 5;
mov ax, 10
mov bx, 20
mul bx
add ax, 5
mov [bp+6], ax
Sample C code
Possible output of a compiler
(using register variables)

Possible output of a compiler
(using code optimization)

One more factor which joins the opposition is that each compiler generates code in its own way. There are very few points where all compilers generate similar code. This means that we need to tailor the decompilation procedure for each compiler, so that if we know which compiler was used to generate this executable, we have a better chance of understanding the code.

For example, the simple "if" statement can be compiled in many ways, like the one given below.

  if (i>20)
{
j=30;
}
else
{
j=40;
}
j++;

    cmp [bp+2], 20
jle lab1

mov [bp+4], 30
jmp lab2
lab1:
mov [bp+4], 40
lab2:
inc [bp+4]

    cmp [bp+2], 20
jg lab1
jmp lab2
lab1:
mov [bp+4], 30
jmp lab3
lab2:
mov [bp+4], 40
lab3:
inc [bp+4]
Sample "if" statement
Possible output #1
Possible output #2

Some other factors which hinder the decompilation process are

1. Self modifying code
Code written for critical applications like games, where every machine cycle counts for the performance, sometimes self-modifying code is used. For example, if the same condition is used at many places in a critical loop, it is evaluated once and the code after that is modified to branch to the same location without evaluating the condition again. But with today's processor speeds, this technique is fast becoming antique!
2. User-defined datatypes
User defined data types, like "struct"s, "typedef"s, "union"s and bit-fields add more to the confusion of the decompiler. Though it is easy to use structures while writing the code, it is almost impossible to figure out if a variable is a part of a structure or is it a basic type on its own, by looking at the compiled output.
3. Use of processor-specific instructions/optimizations

Inspite of all these reasons, it is still possible to decompile a binary executable (though not 100% automated, and not 100% accurate!). Carry on reading...

How to decompile?

A simple approach to decompile a binary executable, is to first parse it and separate it into functions (C style). Once we know where the entry point of the program ("main" for a C program), we can start decompiling that function, and any other function it calls. This way, we can focus on one function at a time carefully.

To use this approach, the first thing that should be known is the entry-point of the program. For a normal C program, it is the "main" function, and for a Win32 program, it is the "WinMain" function. But to find out where these functions begin, we must analyze the executable and figure out which compiler was used, because each compiler has its own entry/exit code added to the program which we need not decompile. If the compiler used is known, we can trace out where the "main" function is starting, and from thereon, till we get a "return" instruction, we can separate out the function.

A simple method to decompile a code snippet

Let's say we have separated out the instructions for a function. To understand how this piece of code works, we need to emulate the processor, and interpret each and every machine instruction! (though it seems a roundabout way, there is nothing called the best method to understand the working of the code, so this is not the only way. If you think this can be done in some other way, please try it out and i'd be interested to hear about it.). And as we are interpreting the code, we need to combine logical group of instructions into simple high level language statements. And voila! We've decompiled the function!

The high level language statements can be grouped as assignment statements, condition evaluation and branches, and function calls. While interpreting the machine code for a function, we have to look for instructions which fall in one of these categories, and accordingly generate the high level code.

Understanding assignment statements and expressions

Take a look at the following code snippet.

   mov ax, [bp+4]
mov bx, 20
mul bx
add ax, 4
mov [bp+4], ax
Sample expression eval & assignment.

It is not difficult to understand what it does. To put it in steps,

  1. AX is loaded with value from memory [bp+4]
  2. BX is loaded with value 20
  3. AX and BX are multiplied, and the result is in DX:AX (of which we will ignore DX for the time being)
  4. A value of 4 is added with AX
  5. The result in AX is stored in memory location [bp+4]

And we can easily write a program which will interpret this code. Let's say we have two variables wAX, wBX. The above 5 steps can be translated by our program as

  1. wAX = [bp+4]
  2. wBX = 20
  3. wAX = wAX*wBX = [bp+4]*20
  4. wAX = wAX+4 = ([bp+4]*20) + 4
  5. [bp+4] = wAX = ([bp+4]*20) + 4

And if we substitute a variable name "i" for "[bp+4]", we have arrived at the high level language statement

i = ( i*20 ) + 4;

That was pretty easy, wasn't it? But then, when do we decide that we have got a full high-level language statement? Whenever there is a instruction which stores some value to external memory, it is a full high-level language statement by itself. Just think about it. All the following high-level language statements inevitably end with a memory store operation!

i = 0;
x = i*30 + (j>>1)+(i/2)
k = (i & 2) + z;

Understanding condition evaluation and branches

A condition evaluation almost always uses a "compare" instruction. So while interpreting the machine code, if we get a "cmp" instruction, we translate it into a "if" statement, and the branch instruction that follows this compare will lead us to the block of code which gets executed if the condition is true, and if it is false.

The following code snippet illustrates this.

   mov ax, [bp+4]
cmp ax, 10
jnz lab1

mov bx, 15
mov [bp+2], bx
jmp lab2
lab1:
mov bx,20
mov [bp+2], bx
lab2:
Sample condition eval & branch.

When we are interpreting this code and come across the "cmp" instruction, the value in our variable wAX is "[bp+4]" (which is a memory variable). So we now form a condition between "[bp+4]" and "10". But what are we comparing for? That can be found using the branch instruction which comes next. Since we find a "jnz" instruction ("jnz" means "jump if not zero", or "jump if not equal"), we can conclude that this is a equality compare i.e. we are comparing for "[bp+4]!=10". The rest is pretty easy, and after replacing "[bp+4]" with a name "i", we can generate the C code for this as below.

if (i != 10) goto lab1;
j = 15;
goto lab2;
lab1:
j = 20;
lab2:

After a bit of rearrangement, and removing the "goto" statements, we can get the 'beautified' C code as

  if (i!= 10)
{
j = 20;
}
else
{
j = 15;
}

But beware! Conditions need not always be evaluated using "compare" instructions. Even arithmetic instructions can set/reset the conditions flag bits in the processor's FLAGS register, and it is these condition flag bits which are actually used to determine if a branch should be taken or not! A classic example is the C statement "i--", as shown below.

 if (i--)
j = 20;
j++;

   sub [bp+4], 1
jz lab1
mov [bp+2], 20
lab1:
add [bp+2], 1
Sample "if" statement
Branch without "compare"

So if we come across such code, we must look at the previous arithmetic instruction, and find out which variable/expression is actually used as the condition.

Interpreting function calls and return values

Function calls are a bit different. The usual "C" convention of function calling, is to push the parameters (last parameter first!) onto the stack, and then call the function. The function will read the values from the stack, and the return value is usually placed in some register (usually the "AX" register). So the trick we can use here is, whenever we find a "push" instruction (which pushes a value on the top of the stack) we make a note of it. And eventually when we reach a "call" instruction (which is actually a subroutine branch instruction), we look at our variables to see what all values/variables have been pushed to the stack, and put them in reverse order to create the high-level C language statement for the function call!

The example for this method is given below. Please note that the code on the right is obtained after replacing "[bp+4]" with "i" and "[bp+2]" with "j".

 mov ax, [bp+4];
push ax
mov ax, [bp+2];
push ax
call _func
mov [bp+4], ax

 
i = func(j, i);
Sample function call
Equivalent C source code.

Sunday, February 15, 2009

AT command set for SONY ERICSSON

The GSM modem will respond in 2 ways, "ERROR" will be returned if the AT command is not supported.If the command is executed successfully an OK will be returned after the response text.(You can use communicate with "GSM" modem via either "Hyperterminal" or cable (Serial port Communication[Java,.NET, Sending and recieving SMS programmatically] ).

e.g AT+CBC=?
response text will be of the form->
+CBC :(0,2),(0-100)
OK


List of AT Commands
--------------------------
  1. AT Attention command
  2. AT* List all supported AT commands
  3. ATZ Restore to user profile (ver. 2)
  4. AT&F Set to factory-defined configuration (ver. 2)
  5. ATI Identification information (ver. 3)
  6. AT&W Store user profile
  7. AT+CLAC List all available AT commands
  8. AT+CGMI Request manufacturer identification (ver. 1)
  9. AT+CGMM Request model identification
  10. AT+CGMR Request revision identification
  11. AT+CGSN Request product serial number identification
  12. AT+GCAP Request modem capabilities list
  13. AT+GMI Request manufacturer information
  14. AT+GMM Request model identification
  15. AT+GMR Request revision identification
  16. ATA Answer incoming call command (ver. 2)
  17. ATH Hook control (ver. 2)
  18. ATD Dial command (ver. 5)
  19. ATO Return to online data mode
  20. AT+CVHU Voice hangup control
  21. AT+CLCC List current calls
  22. AT*CPI Call progress information
  23. ATE Command echo (ver. 2)
  24. ATSO Automatic answer control
  25. ATS2 Escape sequence character
  26. ATS3 Command line termination character (ver. 3)
  27. ATS4 Response formatting character (ver. 3)
  28. ATS5 Command line editing character (ver. 3)
  29. ATS7 Completion connection timeout
  30. ATS10 Automatic disconnect delay control
  31. ATQ Result code suppression (ver. 2)
  32. ATV DCE response mode (ver. 2)
  33. ATX Call progress monitoring control
  34. AT&C Circuit 109 (DCD) control
  35. AT&D Circuit 108 (DTR) response
  36. AT+IFC Cable interface DTE-DCE local flow control
  37. AT+ICF Cable interface character format (ver. 2)
  38. AT+IPR Cable interface port rate
  39. AT+ILRR Cable interface local rate reporting
  40. AT+DS Data compression (ver. 3)
  41. AT+DR Data compression reporting
  42. AT+WS46 Mode selection
  43. AT+FCLASS Select mode
  44. AT*ECBP CHF button pushed (ver. 2)
  45. AT+CMUX Switch to 07.10 multiplexer (ver. 2)
  46. AT*EINA Ericsson system interface active
  47. AT*SEAM Add menu item
  48. AT*SESAF SEMC show and focus
  49. AT*SELERT SEMC create alert (information text)
  50. AT*SESTRI SEMC create string Input
  51. AT*SELIST SEMC create list
  52. AT*SETICK SEMC create ticker
  53. AT*SEDATE SEMC create date field
  54. AT*SEGAUGE SEMC create gauge (bar graph/progress feedback)
  55. AT*SEGUP SEMC update gauge (bar graph/ progress feedback)
  56. AT*SEONO SEMC create on/off input
  57. AT*SEYNQ SEMC create yes/no question
  58. AT*SEDEL SEMC GUI delete
  59. AT*SESLE SEMC soft key label (ver. 1)
  60. AT*SERSK SEMC remove soft key
  61. AT*SEUIS SEMC UI session establish/terminate
  62. AT*EIBA Ericsson Internal Bluetooth address
  63. AT+BINP Bluetooth input
  64. AT+BLDN Bluetooth last dialled number
  65. AT+BVRA Bluetooth voice recognition activation
  66. AT+NREC Noise reduction and echo cancelling
  67. AT+VGM Gain of microphone
  68. AT+VGS Gain of speaker
  69. AT+BRSF Bluetooth retrieve supported
  70. AT+GCLIP Graphical caller ID presentation
  71. AT+CSCS Select TE character set (ver. 3)
  72. AT+CHUP Hang up call
  73. AT+CRC Cellular result codes (ver. 2)
  74. AT+CR Service reporting control
  75. AT+CV120 V.120 rate adaption protocol
  76. AT+VTS DTMF and tone generation
  77. AT+CBST Select bearer service type (ver. 3)
  78. AT+CRLP Radio link protocol (ver. 2)
  79. AT+CEER Extended error report (ver. 2)
  80. AT+CHSD HSCSD device parameters (ver. 2)
  81. AT+CHSN HSCSD non-transparent call configuration (ver. 2)
  82. AT+CHSC HSCSD current call parameters (ver. 2)
  83. AT+CHSR HSCSD parameters report (ver. 2)
  84. AT+CHSU HSCSD automatic user-initiated upgrade
  85. AT+CNUM Subscriber number (ver. 2)
  86. AT+CREG Network registration (ver. 2)
  87. AT+COPS Operator selection (ver. 2)
  88. AT+CLIP Calling line identification (ver. 2)
  89. AT+CLIR Calling line identification restriction
  90. AT+CCFC Calling forwarding number and conditions (ver. 2)
  91. AT+CCWA Call waiting (ver. 2)
  92. AT+CHLD Call hold and multiparty (ver. 1)
  93. AT+CSSN Supplementary service notification (ver. 2)
  94. AT+CAOC Advice of charge
  95. AT+CACM Accumulated call meter (ver. 2)
  96. AT+CAMM Accumulated call meter maximum
  97. AT+CDIP Called line identification presentation
  98. AT+COLP Connected line identification presentation
  99. AT+CPOL Preferred operator list
  100. AT+COPN Read operator names
  101. AT*EDIF Divert function (ver. 2)
  102. AT*EIPS Identify presentation set
  103. AT+CUSD Unstructured supplementary service data (ver. 2)
  104. AT+CLCK Facility lock (ver. 5)
  105. AT+CPWD Change password (Ver. 3)
  106. AT+CFUN Set phone functionality (ver. 2)
  107. AT+CPAS Phone activity status (ver. 3)
  108. AT+CPIN PIN control (ver. 2)
  109. AT+CBC Battery charge (ver. 2)
  110. AT+CSQ Signal quality (ver.1)
  111. AT+CKPD Keypad control (ver. 7)
  112. AT+CIND Indicator control (ver. 5)
  113. AT+CMAR Master reset
  114. AT+CMER Mobile equipment event reporting
  115. AT*ECAM Ericsson call monitoring (ver. 2)
  116. AT+CLAN Language
  117. AT*EJAVA Ericsson Java application function
  118. AT+CSIL Silence Command
  119. AT*ESKL Key-lock mode
  120. AT*ESKS Key sound
  121. AT*EAPP Application function (ver. 5)
  122. AT+CMEC Mobile equipment control mode
  123. AT+CRSM Restricted SIM access
  124. AT*EKSE Ericsson keystroke send
  125. AT+CRSL Ringer sound level (ver. 2)
  126. AT+CLVL Loudspeaker volume level
  127. AT+CMUT Mute control
  128. AT*EMEM Ericsson memory management
  129. AT+CRMP Ring melody playback (ver. 2)
  130. AT*EKEY Keypad/joystick control (ver. 2)
  131. AT*ECDF Ericsson change dedicated file
  132. AT*STKC SIM application toolkit configuration
  133. AT*STKE SIM application toolkit envelope command send
  134. AT*STKR SIM application toolkit command response
  135. AT+CMEE Report mobile equipment error
  136. AT+CSMS Select message service (ver.2)
  137. AT+CPMS Preferred message storage (ver. 4)
  138. AT+CMGF Message format (ver. 1)
  139. AT+CSCA Service centre address (ver. 2)
  140. AT+CSAS Save settings
  141. AT+CRES Restore settings
  142. AT+CNMI New messages indication to TE (ver. 4)
  143. AT+CMGL List message (ver. 2)
  144. AT+CMGR Read message (ver. 2)
  145. AT+CMGS Send message (ver. 2)
  146. AT+CMSS Send from storage (ver. 2)
  147. AT+CMGW Write message to memory (ver. 2)
  148. AT+CMGD Delete message
  149. AT+CMGC Send command (ver. 1)
  150. AT+CMMS More messages to send
  151. AT+CGDCONT Define PDP context (ver. 1)
  152. AT+CGSMS Select service for MO SMS messages
  153. AT+CGATT Packet service attach or detach
  154. AT+CGACT PDP context activate or deactivate
  155. AT+CGDATA Enter data state
  156. AT+CGEREP Packet domain event reporting (ver. 1)
  157. AT+CGREG Packet domain network registration status
  158. AT+CGPADDR Show PDP address
  159. AT+CGDSCONT Define secondary PDP context
  160. AT+CGTFT Traffic flow template
  161. AT+CGEQREQ 3G quality of service profile (requested)
  162. AT+CGEQMIN 3G quality of service profile (minimum acceptable)
  163. AT+CGEQNEG 3G quality of service profile (negotiated)
  164. AT+CGCMOD PDP context modify
  165. Extension of ATD – Request GPRS service
  166. Extension of ATD – Request packet domain IP service
  167. AT+CPBS Phonebook storage (ver. 3)
  168. AT+CPBR Phonebook read (ver. 2)
  169. AT+CPBF Phonebook find (ver. 2)
  170. AT+CPBW Phonebook write (ver. 4)
  171. AT+CCLK Clock (ver. 4)
  172. AT+CALA Alarm (ver. 3)
  173. AT+CALD Alarm delete
  174. AT+CAPD Postpone or dismiss an alarm (ver. 2)
  175. AT*EDST Ericsson daylight saving time
  176. AT+CIMI Request international mobile subscriber identity
  177. AT*EPEE PIN event
  178. AT*EAPS Active profile set
  179. AT*EAPN Active profile rename
  180. AT*EBCA Battery and charging algorithm (ver. 4)
  181. AT*ELIB Ericsson list Bluetooth devices
  182. AT*EVAA Voice answer active (ver. 1)
  183. AT*EMWS Magic word set
  184. AT+CPROT Enter protocol mode
  185. AT*EWDT WAP download timeout
  186. AT*EWBA WAP bookmark add (ver. 2)
  187. AT*EWCT WAP connection timeout
  188. AT*EIAC Internet account, create
  189. AT*EIAD Internet account configuration, delete
  190. AT*EIAW Internet account configuration, write general parameters
  191. AT*EIAR Internet account configuration, read general parameters
  192. AT*EIAPSW Internet account configuration, write PS bearer parameters
  193. AT*EIAPSR Internet account configuration, read PS bearer parameters
  194. AT*EIAPSSW Internet account configuration, write secondary PDP context parameters
  195. AT*EIAPSSR Internet account configuration, read secondary PDP context parameters
  196. AT*EIACSW Internet account configuration, write CSD bearer parameters
  197. AT*EIACSR Internet account configuration, read CSD bearer parameters
  198. AT*EIABTW Internet account configuration, write Bluetooth bearer parameters
  199. AT*EIABTR Internet account configuration, read Bluetooth bearer parameters
  200. AT*EIAAUW Internet account configuration, write authentication parameters
  201. AT*EIAAUR Internet account configuration, read authentication parameters
  202. AT*EIALCPW Internet account configuration, write PPP parameters – LCP
  203. AT*EIALCPR Internet account configuration, read PPP parameters – LCP
  204. AT*EIAIPCPW Internet account configuration, write PPP parameters – IPCP
  205. AT*EIAIPCPR Internet account configuration, read PPP parameters – IPCP
  206. AT*EIADNSV6W Internet account configuration, write DNS parameters – IPv6CP
  207. AT*EIADNSV6R Internet account configuration, read DNS parameters – IPv6CP
  208. AT*EIARUTW Internet account configuration, write routing table parameters
  209. AT*EIARUTD Internet account configuration, delete routing table parameters
  210. AT*EIARUTR Internet account configuration, read routing table parameters
  211. AT*SEACC Accessory class report
  212. AT*SEACID Accessory identification
  213. AT*SEACID2 Accessory identification (Bluetooth)
  214. AT*SEAUDIO Accessory class report
  215. AT*SECHA Charging control
  216. AT*SELOG SE read log
  217. AT*SEPING SE ping command
  218. AT*SEAULS SE audio line status
  219. AT*SEFUNC SE functionality status (ver. 2)
  220. AT*SEFIN SE flash Information
  221. AT*SEFEXP Flash auto exposure setting from ME
  222. AT*SEMOD Camera mode indicator to the flash
  223. AT*SEREDI Red eye reduction indicator to the flash
  224. AT*SEFRY Ready indicator to the ME
  225. AT*SEAUP Sony Ericsson audio parameters
  226. AT*SEVOL Volume level
  227. AT*SEVOLIR Volume indication request
  228. AT*SEBIC Status bar icon
  229. AT*SEANT Antenna identification
  230. AT*SESP Speakermode on/off
  231. AT*SETBC Text to bitmap converter
  232. AT*SEAVRC Sony Ericsson audio video remote control
  233. AT*SEMMIR Sony Ericsson multimedia information request
  234. AT*SEAPP Sony Ericsson application
  235. AT*SEAPPIR Sony Ericsson application indication request
  236. AT*SEJCOMM Sony Ericsson Java comm
  237. AT*SEDUC Sony Ericsson disable USB charge
  238. AT*SEABS Sony Ericsson accessory battery status
  239. AT*SEAVRCIR Sony Ericsson audio video remote control indication request
  240. AT*SEGPSA Sony Ericsson global positioning system accessory
  241. AT*SEAUDIO Accessory class report
  242. AT*SEGPSA Sony Ericsson global positioning system accessory
  243. AT*SEAUDIO Accessory Class Report
  244. AT*SEGPSA Sony Ericsson global positioning system accessory
  245. AT*SETIR Sony Ericsson time information request
  246. AT*SEMCM Sony Ericsson memory card management
  247. AT*SEAUDIO Accessory Class Report

Saturday, February 7, 2009

Top 10 hacking incidents.

Top 10 hacking incidents of all time instances where some of the most seemingly secure computer networks were compromised.


Early 1990s :
Kevin Mitnick, often incorrectly called by many as god of hackers, broke into the computer systems of the world's top technology and telecommunications companies Nokia, Fujitsu, Motorola, and Sun Microsystems. He was arrested by the FBI in 1995, but later released on parole in 2000. He never termed his activity hacking, instead he called it social engineering.

November 2002:
Englishman Gary McKinnon was arrested in November 2002 following an accusation that he hacked into more than 90 US military computer systems in the UK. He is currently undergoing trial in a British court for a "fast-track extradition" to the US where he is a wanted man. The next hearing in the case is slated for today.

1995:
Russian computer geek Vladimir Levin effected what can easily be called The Italian Job online - he was the first person to hack into a bank to extract money. Early 1995, he hacked into Citibank and robbed $10 million. Interpol arrested him in the UK in 1995, after he had transferred money to his accounts in the US, Finland, Holland, Germany and Israel.

1990:
When a Los Angeles area radio station announced a contest that awarded a Porsche 944S2 for the 102nd caller, Kevin Poulsen took control of the entire city's telephone network, ensured he is the 102nd caller, and took away the Porsche beauty. He was arrested later that year and sentenced to three years in prison. He is currently a senior editor at Wired News.

1983:
Kevin Poulsen again. A little-known incident when Poulsen, then just a student, hacked into Arpanet, the precursor to the Internet was hacked into. Arpanet was a global network of computers, and Poulsen took advantage of a loophole in its architecture to gain temporary control of the US-wide network.

1996:
US hacker Timothy Lloyd planted six lines of malicious software code in the computer network of Omega Engineering which was a prime supplier of components for NASA and the US Navy. The code allowed a "logic bomb" to explode that deleted software running Omega's manufacturing operations. Omega lost $10 million due to the attack.

1988:
Twenty-three-year-old Cornell University graduate Robert Morris unleashed the first Internet worm on to the world. Morris released 99 lines of code to the internet as an experiment, but realised that his program infected machines as it went along. Computers crashed across the US and elsewhere. He was arrested and sentenced in 1990.

1999:
The Melissa virus was the first of its kind to wreak damage on a global scale. Written by David Smith (then 30), Melissa spread to more than 300 companies across the world completely destroying their computer networks. Damages reported amounted to nearly $400 million. Smith was arrested and sentenced to five years in prison.

2000:
MafiaBoy, whose real identity has been kept under wraps because he is a minor, hacked into some of the largest sites in the world, including eBay, Amazon and Yahoo between February 6 and Valentine's Day in 2000. He gained access to 75 computers in 52 networks, and ordered a Denial of Service attack on them. He was arrested in 2000.

1993:
They called themselves Masters of Deception, targeting US phone systems. The group hacked into the National Security Agency, AT&T, and Bank of America. It created a system that let them bypass long-distance phone call systems, and gain access to private lines.

Friday, February 6, 2009

Web Crawler or WebRobot or Web Spider Working

A web spider, some times called a crawler or a robot, plays an important role as an essential infrastructure of every search engines. It automatically discovers and collects resources, especially the web pages, from the Internet. As the rapidly growth of the Internet, a typical design of web spider may not cope with the overwhelming number of web pages.

Search engines.
A search engine is a program that searches through some dataset. In the context of the Web, the word "search engine" is most often used for search forms that search through databases of HTML documents gathered by a robot .Robots are software agents.

Web Agent
The word "agent" is used for lots of meanings in computing these days. Specifically:
Autonomous agents are programs that do travel between sites, deciding themselves when to move and what to do. These can only travel between special servers and are currently not widespread in the Internet. Intelligent agents are programs that help users with things, such as choosing a product, or guiding a user through form filling, or even helping users find things. These have generally little to do with networking. User-agent is a technical name for programs that perform networking tasks for a user, such as Web User-agents like Netscape Navigator and Microsoft Internet Explorer, and Email User-agent like Qualcomm Eudora etc.

This process is called web crawling or spidering. Many sites, in particular search engines, use spidering as a means of providing up-to-date data. Web crawlers are mainly used to create a copy of all the visited pages for later processing by a search engine that will index the downloaded pages to provide fast searches. Crawlers can also be used for automating maintenance tasks on a website, such as checking links or validating HTML code. Also, crawlers can be used to gather specific types of information from Web pages, such as harvesting e-mail addresses (usually for spam).

A web crawler is one type of bot, or software agent. In general, it starts with a list of URLs to visit, called the seeds. As the crawler visits these URLs, it identifies all the hyperlinks in the page and adds them to the list of URLs to visit, called the crawl frontier. URLs from the frontier are recursively visited according to a set of policies.

A robot is a program that automatically traverses the Web's hypertext structure by retrieving a document, and recursively retrieving all documents that are referenced. Note that "recursive" here doesn't limit the definition to any specific traversal algorithm; even if a robot applies some heuristic to the selection and order of documents to visit and spaces out requests over a long space of time, it is still a robot. Normal Web browsers are not robots, because the are operated by a human, and don't automatically retrieve referenced documents (other than inline images).Web robots are sometimes referred to as Web Wanderers, Web Crawlers, or Spiders. These names are a bit misleading as they give the impression the software itself moves between sites like a virus; this not the case, a robot simply visits sites by requesting documents from them.

Basic Search engine Architecture


searchengine+architecture.bmp
Before a search engine can tell you where a file or document is, it must be found. To find information on the hundreds of millions of Web pages that exist, a typical search engine employs special software robots, called spiders, to build lists of the words found on Web sites. When a spider is building its lists, the process is called Web crawling. A Web crawler is a program, which automatically traverses the web by downloading documents and following links from page to page. They are mainly used by web search engines to gather data for indexing. Other possible applications include page validation, structural analysis and visualization, update notification, mirroring and personal web assistants/agents etc. Web crawlers are also known as spiders, robots, worms etc.



Crawlers are automated programs that follow the links found on the web pages.


There is a URL Server that sends lists of URLs to be fetched to the crawlers. The web pages that are fetched are then sent to the store server. The store server then compresses and stores the web pages into a repository. Every web page has an associated ID number called a doc ID, which is assigned whenever a new URL is parsed out of a web page. The indexer and the sorter perform the indexing function. The indexer performs a number of functions. It reads the repository, uncompresses the documents, and parses them. Each document is converted into a set of word occurrences called hits. The hits record the word, position in document, an approximation of font size, and capitalization. The indexer distributes these hits into a set of "barrels", creating a partially sorted forward index. The indexer performs another important function. It parses out all the links in every web page and stores important information about them in an anchors file. This file contains enough information to determine where each link points from and to, and the text of the link.


The URL Resolver reads the anchors file and converts relative URLs into absolute URLs and in turn into doc IDs. It puts the anchor text into the forward index, associated with the doc ID that the anchor points to. It also generates a database of links, which are pairs of doc IDs. The links database is used to compute Page Ranks for all the documents. The sorter takes the barrels, which are sorted by doc ID and resorts them by word ID to generate the inverted index. This is done in place so that little temporary space is needed for this operation. The sorter also produces a list of word IDs and offsets into the inverted index.


A program called Dump Lexicon takes this list together with the lexicon produced by the indexer and generates a new lexicon to be used by the searcher. A lexicon lists all the terms occurring in the index along with some term-level statistics (e.g., total number of documents in which a term occurs) that are used by the ranking algorithms The searcher is run by a web server and uses the lexicon built by Dump Lexicon together with the inverted index and the Page Ranks to answer queries. (Brin and Page 1998).

Search Engine Architecture

How a Web Crawler Works



Web crawlers are an essential component to search engines; running a web crawler is a challenging task. There are tricky performance and reliability issues and even more importantly, there are social issues. Crawling is the most fragile application since it involves interacting with hundreds of thousands of web servers and various name servers, which are all beyond the control of the system. Web crawling speed is governed not only by the speed of one’s own Internet connection, but also by the speed of the sites that are to be crawled. Especially if one is a crawling site from multiple servers, the total crawling time can be significantly reduced, if many downloads are done in parallel. Despite the numerous applications for Web crawlers, at the core they are all fundamentally the same. Following is the process by which Web crawlers work:

1. Download the Web page.

2. Parse through the downloaded page and retrieve all the links.

3. For each link retrieved, repeat the process

Architecture of web crawler


Web Crawler Architecture


The Web crawler can be used for crawling through a whole site on the Inter-/Intranet. You specify a start-URL and the Crawler follows all links found in that HTML page. This usually leads to more links, which will be followed again, and so on. A site can be seen as a tree-structure, the root is the start-URL; all links in that root-HTML-page are direct sons of the root. Subsequent links are then sons of the previous sons.


A single URL Server serves lists of URLs to a number of crawlers. Web crawler starts by parsing a specified web page, noting any hypertext links on that page that point to other web pages. They then parse those pages for new links, and so on, recursively.


Webcrawler software doesn't actually move around to different computers on the Internet, as viruses or intelligent agents do. Each crawler keeps roughly 300 connections open at once. This is necessary to retrieve web pages at a fast enough pace. A crawler resides on a single machine. The crawler simply sends HTTP requests for documents to other machines on the Internet, just as a web browser does when the user clicks on links. All the crawler really does is to automate the process of following links.


Web crawling can be regarded as processing items in a queue. When the crawler visits a web page, it extracts links to other web pages. So the crawler puts these URLs at the end of a queue, and continues crawling to a URL that it removes from the front of the queue.

Crawling policies

There are three important characteristics of the Web that generate a scenario in which Web crawling is very difficult:

· its large volume,

· its fast rate of change, and

· dynamic page generation,

which combine to produce a wide variety of possible crawlable URLs.


The large volume implies that the crawler can only download a fraction of the web pages within a given time, so it needs to prioritize its downloads. The high rate of change implies that by the time the crawler is downloading the last pages from a site, it is very likely that new pages have been added to the site, or that pages have already been updated or even deleted.


The recent increase in the number of pages being generated by server-side scripting languages has also created difficulty in that endless combinations of HTTP GET parameters exist, only a small selection of which will actually return unique content. For example, a simple online photo gallery may offer three options to users, as specified through HTTP GET parameters. If there exist four ways to sort images, three choices of thumbnail size, two file formats, and an option to disable user-provided contents, then that same set of content can be accessed with forty-eight different URLs, all of which will be present on the site. This mathematical combination creates a problem for crawlers, as they must sort through endless combinations of relatively minor scripted changes in order to retrieve unique content.

The behavior of a web crawler is the outcome of a combination of policies:

· A selection policy that states which pages to download.

· A re-visit policy that states when to check for changes to the pages.

· A politeness policy that states how to avoid overloading websites.

· A parallelization policy that states how to coordinate distributed web crawlers

Selection policy


Given the current size of the Web, even large search engines cover only a portion of the publicly available internet; a study by Lawrence and Giles (Lawrence and Giles, 2000) showed that no search engine indexes more than 16% of the Web. As a crawler always downloads just a fraction of the Web pages, it is highly desirable that the downloaded fraction contains the most relevant pages, and not just a random sample of the Web.


This requires a metric of importance for prioritizing Web pages. The importance of a page is a function of its intrinsic quality, its popularity in terms of links or visits, and even of its URL (the latter is the case of vertical search engines restricted to a single top-level domain, or search engines restricted to a fixed Web site). Designing a good selection policy has an added difficulty: it must work with partial information, as the complete set of Web pages is not known during crawling.



Different types of crawling.

Path-ascending crawling


Some crawlers intend to download as many resources as possible from a particular Web site. Cothey introduced a path-ascending crawler that would ascend to every path in each URL that it intends to crawl.

Focused crawling
The importance of a page for a crawler can also be expressed as a function of the similarity of a page to a given query. Web crawlers that attempt to download pages that are similar to each other are called focused crawler or topical crawlers.



Crawling the Deep Web


A vast amount of Web pages lie in the deep or invisible Web. These pages are typically only accessible by submitting queries to a database, and regular crawlers are unable to find these pages if there are no links that point to them. Google’s Sitemap Protocol and mod oai (Nelson et al., 2005) are intended to allow discovery of these deep-Web resources.

Re-visit policy


The Web has a very dynamic nature, and crawling a fraction of the Web can take a really long time, usually measured in weeks or months. By the time a web crawler has finished its crawl, many events could have happened. These events can include creations, updates and deletions.

Uniform policy: This involves re-visiting all pages in the collection with the same frequency, regardless of their rates of change.

Proportional policy: This involves re-visiting more often the pages that change more frequently. The visiting frequency is directly proportional to the (estimated) change frequency.

Politeness policy

Crawlers can retrieve data much quicker and in greater depth than human searchers, so they can have a crippling impact on the performance of a site. Needless to say if a single crawler is performing multiple requests per second and/or downloading large files, a server would have a hard time keeping up with requests from multiple crawlers.

As noted by Koster (Koster, 1995), the use of web crawlers is useful for a number of tasks, but comes with a price for the general community. The costs of using web crawlers include:

Network resources, as crawlers require considerable bandwidth and operate with a high degree of parallelism during a long period of time.

Server overload, especially if the frequency of accesses to a given server is too high.

Poorly written crawlers, which can crash servers or routers, or which download pages they cannot handle.

Personal crawlers that, if deployed by too many users, can disrupt networks and Web servers.

Parallelization policy

A parallel crawler is a crawler that runs multiple processes in parallel. The goal is to maximize the download rate while minimizing the overhead from parallelization and to avoid repeated downloads of the same page. To avoid downloading the same page more than once, the crawling system requires a policy for assigning the new URLs discovered during the crawling process, as the same URL can be found by two different crawling processes.

Crawling is an effective process synchronisation tool between the users and the search engine.

Web Robot Algorithms


Each robot uses different algorithms to decide where to visit. In general, they start from a historical list of URLS, especially some most popular web sites on the Web.




Starting at a location on the web reveals a branching structure which, if cycles are avoided, is essentially a tree

· Depth First Traversal

· Breadth First Traversal

· Heuristics search

For each URL(web page), we use a heuristics function to evaluate its importance. Then we visit those important web pages first

Robots Exclusion Standard


The robots exclusion standard, also known as the Robots Exclusion Protocol or robots.txt protocol is a convention to prevent cooperating web spiders and other web robots from accessing all or part of a website which is, otherwise, publicly viewable. Robots are often used by search engines to categorize and archive web sites, or by webmasters to proofread source code. The standard complements Sitemaps, a robot inclusion standard for websites.


A robots.txt file on a website will function as a request that specified robots ignore specified files or directories in their search. This might be, for example, out of a preference for privacy from search engine results, or the belief that the content of the selected directories might be misleading or irrelevant to the categorization of the site as a whole, or out of a desire that an application only operate on certain data



This example allows all robots to visit all files because the wildcard "*" specifies all robots:

User-agent: *

Disallow:

This example keeps google robots out:

User-agent: googlebot

Disallow: /

The next is an example that tells all crawlers not to enter into four directories of a website:

User-agent: *

Disallow: /cgi-bin/

Disallow: /images/

Disallow: /tmp/

Disallow: /private/

Example that tells a specific crawler not to enter one specific directory:

User-agent: BadBot

Disallow: /private/


Will the /robots.txt standard be extended?

Probably... there are some ideas floating around. They haven't made it into a coherent proposal because of time constraints, and because there is little pressure. Mail suggestions to the robots mailing list, and check the robots home page for work in progress.



What if I can't make a /robots.txt file?



Sometimes you cannot make a /robots.txt file, because you don't administer the entire server. All is not lost: there is a new standard for using HTML META tags to keep robots out of your documents.

Googlebot, Google’s Web Crawler

Googlebot is Google’s web crawling robot, which finds and retrieves pages on the web and hands them off to the Google indexer. It’s easy to imagine Googlebot as a little spider scurrying across the strands of cyberspace, but in reality Googlebot doesn’t traverse the web at all. It functions much like your web browser, by sending a request to a web server for a web page, downloading the entire page, then handing it off to Google’s indexer.

Googlebot consists of many computers requesting and fetching pages much more quickly than you can with your web browser. In fact, Googlebot can request thousands of different pages simultaneously. To avoid overwhelming web servers, or crowding out requests from human users, Googlebot deliberately makes requests of each individual web server more slowly than it’s capable of doing.



Google’s Indexer

Googlebot gives the indexer the full text of the pages it finds. These pages are stored in Google’s index database. This index is sorted alphabetically by search term, with each index entry storing a list of documents in which the term appears and the location within the text where it occurs. This data structure allows rapid access to documents that contain user query terms.

To improve search performance, Google ignores (doesn’t index) common words called stop words (such as the, is, on, or, of, how, why, as well as certain single digits and single letters). Stop words are so common that they do little to narrow a search, and therefore they can safely be discarded. The indexer also ignores some punctuation and multiple spaces, as well as converting all letters to lowercase, to improve Google’s performance.

Google’s Query Processor

The query processor has several parts, including the user interface (search box), the “engine” that evaluates queries and matches them to relevant documents, and the results formatter.

PageRank is Google’s system for ranking web pages. A page with a higher PageRank is deemed more important and is more likely to be listed above a page with a lower PageRank.

Google considers over a hundred factors in computing a PageRank and determining which documents are most relevant to a query, including the popularity of the page, the position and size of the search terms within the page, and the proximity of the search terms to one another on the page. A patent application discusses other factors that Google considers when ranking a page. Visit SEOmoz.org’s report for an interpretation of the concepts and the practical applications contained in Google’s patent application.

Google also applies machine-learning techniques to improve its performance automatically by learning relationships and associations within the stored data. For example, the spelling-correcting system uses such techniques to figure out likely alternative spellings. Google closely guards the formulas it uses to calculate relevance; they’re tweaked to improve quality and performance, and to outwit the latest devious techniques used by spammers.

Indexing the full text of the web allows Google to go beyond simply matching single search terms. Google gives more priority to pages that have search terms near each other and in the same order as the query. Google can also match multi-word phrases and sentences. Since Google indexes HTML code in addition to the text on the page, users can restrict searches on the basis of where query words appear, e.g., in the title, in the URL, in the body, and in links to the page, options offered by Google’s Advanced Search Form and Using Search Operators (Advanced Operators).
Let’s see how Google processes a query.

How google query traverse



Demerits of web crawler

It requires considerable bandwidth.

It sometime uses of spamming.

Unable to crawl all deep web.

Dynamic content - dynamic pages which are returned in response to a submitted query or accessed only through a form (especially if open-domain input elements e.g. text fields are used; such fields are hard to navigate without domain knowledge).

Unlinked content - pages which are not linked to by other pages, which may prevent Web crawling programs from accessing the content. This content is referred to as pages without backlinks (or inlinks).

Private Web - sites that require registration and login (password-protected resources).

Contextual Web - pages with content varying for different access contexts (e.g. ranges of client IP addresses or previous navigation sequence).

Limited access content - sites that limit access to their pages in a technical way (e.g., using the Robots Exclusion Standard, CAPTCHAs or pragma:no-cache/cache-control:no-cache HTTP headers), prohibiting search engines from browsing them and creating cached copies.

Scripted content - pages that are only accessible through links produced by JavaScript as well as content dynamically downloaded from Web servers via Flash or AJAX solutions.

Non-HTML/text content - textual content encoded in multimedia (image or video) files or specific file formats not handled by search engines.



  • Poorly written web robots may damage files in server.

  • Certain robot implementations can (and have in the past) overloaded networks and servers. This happens especially with people who are just starting to write a robot; these days there is sufficient information on robots to prevent some of these mistakes.

  • Robots are operated by humans, who make mistakes in configuration, or simply don't consider the implications of their actions. This means people need to be careful, and robot authors need to make it difficult for people to make mistakes with bad effects

  • · Web-wide indexing robots build a central database of documents, which doesn't scale too well to millions of documents on millions of sites



Guidelines for robot writers

To write a good Web Robot, you should try to avoid

· Overloading network

· Overload a server with rapid requests for documents

· Servers that unreachable

· Cycles in the web structure

Also

Be Accountable

Test Locally

Stay with it

Don't hog resources

Share results , If you are interested in writing your own crawler Please comment on my post with (which language like C#, java) language.We will give you assistance through our blog.Also if you want clarification for any of area of our post please feel free to comment.

webdevlopers  webhosting company cochin/kerala/india

Thanks.


References:

http://www.google.com/technology/

http://www.robotstxt.org/wc/robots.html

http://www.robotstxt.org/wc/exclusion.html

What is Web analytics or Web Matrices?

Web analytics is the process of collecting data about the activities of people accessing your website (i.e. visitors), how they found you, when they visited, what pages they looked at, what they bought or downloaded and so on, and mining that data for information that can be used to improve your website. In general, Web analytics is the detailed statistics about the visitors to a website.

Also we can say that Web analytics is the study of the behavior of website visitors. In a commercial context, web analytics especially refers to the use of data collected from a web site to determine which aspects of the website work towards the business objectives; for example, which landing pages encourage people to make a purchase.


Users of Web analytics can define and track conversions, or goals. Goals might include sales, lead generation, viewing of a specific page, or download a particular file. By using this tool, marketers can determine which ads are performing, and which are not, as well as find unexpected sources of quality visitors.

There are two main technological approaches to collecting web analytics data. The first method called logfile analysis reads the log files in which the web server records all its transactions. The second method called page tagging, uses JavaScript on each page to notify a third-party server when a web browser renders a page or it may be hybrid type. If you want to know more about web analytics and creating your own web analytic system Please visit our blog again our next article on web analytics give you idea about how to create a web analytics system i.e Website visitor tracking.
FREE Web analytics Service by Google:www.google.com/analytics

by,
cochin webdevelopers seoservices kerala

LinkWithin

Related Posts with Thumbnails