Showing posts with label Ragh Juluri. Show all posts
Showing posts with label Ragh Juluri. Show all posts

Thursday, November 07, 2013

Steps to Configure Google Talk Account in Pidgin



If you’ve ever tried to setup your Google Talk account for your own domain in the Pidgin multi-protocol instant messenger client. Here’s how to do it.
Open up Pidgin and choose Accounts –> Manage Accounts.

image

Then click the Add button.
Then you’ll want to enter in your username as the beginning of your email address, and the Domain as the part after the @ symbol. For instance, mine is abc@gmail.com, so I’m using abc as the username and gmail.com as the Domain.
Then flip over to the Advanced tab, and Connection Security : Requires encrypiton and enter talk.google.com as the Connect server. You can pretty much leave the other settings alone if you want.
And flip to Proxy tab and enter proxy details
Proxy Type: HTTP
Host : www-proxy.xyz.com
Port : 80 

Click on Save Button and restart Pidgin. Now all google talk contacts will be visible in pidgin and able to initiate conversation.

Thursday, March 28, 2013

Benchmarking a Hadoop Cluster




Benchmarks make good tests because you also get numbers that you can compare with other clusters as a sanity check on whether your new cluster is performing roughly as expected. And you can tune a cluster using benchmark results to squeeze the best performance out of it.


To get the best results, you should run benchmarks on a cluster that is not being used by others. In practice, this is just before it is put into service and users start relying on it. Once users have scheduled periodic jobs on a cluster, it is generally impossible to find a time when the cluster is not being used.


Hadoop comes with several benchmarks that you can run very easily with minimal setup cost. Benchmarks are packaged in the test JAR file, and you can get a list of them, with descriptions, by invoking the JAR file with no arguments:

% hadoop jar $HADOOP_INSTALL/hadoop-*-test.jar

Most of the benchmarks show usage instructions when invoked with no arguments.

For example:

% hadoop jar $HADOOP_INSTALL/hadoop-*-test.jar TestDFSIO 

TestFDSIO.0.0.4
Usage: TestFDSIO -read | -write | -clean [-nrFiles N] [-fileSize MB] [-resFile
resultFileName] [-bufferSize Bytes]


Benchmarking HDFS with TestDFSIO

TestDFSIO tests the I/O performance of HDFS. It does this by using a MapReduce job as a convenient way to read or write files in parallel. Each file is read or written in a separate map task, and the output of the map is used for collecting statistics related to the file just processed. The statistics are accumulated in the reduce to produce a summary.

The following command writes 10 files of 1,000 MB each:

% hadoop jar $HADOOP_INSTALL/hadoop-*-test.jar TestDFSIO -write -nrFiles 10
-fileSize 1000


At the end of the run, the results are written to the console and also recorded in a local
file (which is appended to, so you can rerun the benchmark and not lose old results):
% cat TestDFSIO_results.log
----- TestDFSIO ----- : write
Date & time: Sun Apr 12 07:14:09 EDT 2009
Number of files: 10
Total MBytes processed: 10000
Throughput mb/sec: 7.796340865378244
Average IO rate mb/sec: 7.8862199783325195
IO rate std deviation: 0.9101254683525547
Test exec time sec: 163.387
The files are written under the /benchmarks/TestDFSIO directory


To run a read benchmark, use the -read argument. Note that these files must already
exist (having been written by TestDFSIO -write):

% hadoop jar $HADOOP_INSTALL/hadoop-*-test.jar TestDFSIO -read -nrFiles 10
-fileSize 1000

Here are the results for a real run:

----- TestDFSIO ----- : read
Date & time: Sun Apr 12 07:24:28 EDT 2009
Number of files: 10
Total MBytes processed: 10000
Throughput mb/sec: 80.25553361904304
Average IO rate mb/sec: 98.6801528930664
IO rate std deviation: 36.63507598174921
Test exec time sec: 47.624
When you’ve finished benchmarking, you can delete all the generated files from HDFS
using the -clean argument:

% hadoop jar $HADOOP_INSTALL/hadoop-*-test.jar TestDFSIO -clean


Benchmarking MapReduce with Sort

Hadoop comes with a MapReduce program that does a partial sort of its input. It is very useful for benchmarking the whole MapReduce system, as the full input dataset is transferred through the shuffle. The three steps are: generate some random data, perform the sort, then validate the results.

First, we generate some random data using RandomWriter. It runs a MapReduce job with 10 maps per node, and each map generates (approximately) 1 GB of random binary data, with keys and values of various sizes.



Here’s how to invoke RandomWriter (found in the example JAR file, not the test one) to write its output to a directory called random-data:

% hadoop jar $HADOOP_INSTALL/hadoop-*-examples.jar randomwriter random-data

Next, we can run the Sort program:


% hadoop jar $HADOOP_INSTALL/hadoop-*-examples.jar sort random-data sorted-data

The overall execution time of the sort is the metric we are interested in, but it’s instructive to watch the job’s progress via the web UI (http://jobtracker-host:50030/), where you can get a feel for how long each phase of the job takes.


sanity check, we validate that the data in sorted-data is, in fact, correctly sorted:

% hadoop jar $HADOOP_INSTALL/hadoop-*-test.jar testmapredsort -sortInput random-data \

-sortOutput sorted-data

This command runs the SortValidator program, which performs a series of checks on the unsorted and sorted data to check whether the sort is accurate. It reports the outcome to the console at the end of its run:
SUCCESS! Validated the MapReduce framework's 'sort' successfully.


MRBench (invoked with mrbench) runs a small job a number of times. It acts as a good
counterpoint to sort, as it checks whether small job runs are responsive.

NNBench (invoked with nnbench) is useful for load-testing namenode hardware.

Gridmix is a suite of benchmarks designed to model a realistic cluster workload by
mimicking a variety of data-access patterns seen in practice. See the documentation
in the distribution for how to run Gridmix








Thursday, July 26, 2012

PROVIDENT FUND (PF) WITHDRAWAL : IBM INDIA PVT LTD

PROVIDENT FUND (PF)  WITHDRAWAL : IBM INDIA PVT LTD

In the Current Blog..i want to describe the guidelines for withdrawing Provident Fund (PF) in IBM India Pvt Ltd. I just wanted to explain the step-by-step procedure to be followed .. as many of my friends asked me the procedure for withdrawal ..



Note: Provident Fund Withdrawal needs to be initiated only after the completion of 60 days from the date of leaving Or else it will be rejected.


Please fill the Forms attached (Form-10C.pdf -Form-19.pdf and Guidelines-form-10C.pdf , Guidelines-Form-19.pdf)  and send the duly filled in and signed  hardcopies of the forms to the address below for withdrawal of PF



Please refer the sample filled form which is attached for your reference,
the details which shown in the sample form is mandatory, if you
miss any of details then your Form will be rejected and sent back to
you.

Please ensure below points are taken care before sending the forms to
IBM : 

1. Application Forms – Please ensure that the Forms are printed back to back.

2. Filling the Application - Please use only blue ink to fill the Forms.

3. Name – Please ensure you are filling your name same as it appears in IBM records. Also ensure the name is matching with your bank account. Any mismatch in this regard will lead to rejection of the claim.

4. Over Writing – Please avoid over writing in the form, as over writing will lead to rejection of the claim.


5. Signatures - Please ensure that your application is complete with all mandatory employee signatures. (3 Signatures on Page No.3 of the Form 19 and 2 signatures on Page No.3 of the Form 10C).


6. Cancelled cheque leaf - Please ensure that a cancelled cheque leaf
(with the name,IFSC code printed on it) has been attached along with
the PF withdrawal form. Also ensure that the bank details filled in
the form is matching with the details appearing in attached cancelled
cheque leaf. In case name is not printed on cheque, along with
cheque leaf ensure that a bank statement with last few transactions
is attached with forms. Please make sure IFSC code is appearing in
both cancelled cheque leaf and bank statement. Also make sure that
the bank account number you are providing is of individual account.


7. ID Proof- In case you are applying for PF withdrawal after 2 years
and 6 months from the date of leaving , ensure that you have
attached Photo ID proofs Like Pan card/ Passport/ Voters ID along
with the forms.



8. Please enclose following documents along with forms
    a) photocopy of Full and Final settlement letter.
    b) Please attach a (crossed cheque leaf written as cancelled).
    c) cancelled cheque.

9. Scanned/ photocopy applications are not acceptable for process.

10.Please mention your mobile number on top of the application (written in
pencil ).

Above documents needs to be couriered to the following address.

IBM India Pvt Ltd.
Retirals Team
Global Process Services - HR Delivery
Manyata Embassy Business Park,
D1, 4th Floor, Outer Ring Road,
Nagawara, Bangalore - 560 045
ibmretirals@aonhewitt.com

Once above forms are couriered to IBM address, they will file with PF office 

After 1 Month (approx) you will receive pf withdrawal tracking number to your mobile ..which you mentioned earlier.

Then you need to wait for 3 months (approx) for the PF amount to be credited in your Bank Account which you mentioned earlier.If you file during march-july months it will take long time as Provident Fund office employees are involved in the year end financial transaction calculations .. it might take up to 6 months .. 

PF Withdrawal will take less time compared to PF Transfer .. some times it will take up to 2 years for PF transfer .. So in my perspective PF Withdrawal is better Option then PF Transfer

Uploaded all 4 documents to Google Docs .. so please download from the following urls








Friday, July 20, 2012

HIVE - HADOOP : INSTALLATION , EXECUTION OF SAMPLE PROGRAM - Part II

HIVE - HADOOP : INSTALLATION , EXECUTION  OF SAMPLE PROGRAM - Part II

Hive Clients

To Start Hive Server use following command .. 

$hive --service hiveserver

Following are different Hive Clients, to connect to this Server

1) CommandLine
2) Thrift Client
3) JDBC Driver
4) ODBC Driver


Courtesy : HadoopDefinativeGuide

1) CommandLine : Operates in embedded mode only, it needs to have access to the hive libraries.

This is section is briefly described in my previous article ( Hive - Part I ) .

2) Thrift Client : 

The Apache Thrift software framework, for scalable cross-language services development, combines a software stack with a code generation engine to build services that work efficiently and seamlessly between C++, Java, Python, PHP, Ruby, Erlang, Perl, Haskell, C#, Cocoa, JavaScript, Node.js, Smalltalk, OCaml and Delphi and other languages. Thrift Software

Hive Server is implemented in java, so to query hive server using ruby,php,C++  language then you need to build thrift client specific to that language.

Steps to build Thrift Client for Ruby

a) Install Thrift Thrift
b)  Download Thrift Source Thrift Source
c) Navigate to Hive Source directory 

thrift —gen rb -I service/include metastore/if/hive_metastore.thrift 
thrift —gen rb -I service/include -I . service/if/hive_service.thrift 
thrift —gen rb service/include/thrift/fb303/if/fb303.thrift 
thrift —gen rb serde/if/serde.thrift 
thrift —gen rb ql/if/queryplan.thrift 
thrift —gen rb service/include/thrift/if/reflection_limited.thrift

Or you can download Thrift client for ruby from Thrift Ruby Client

Thrift Java Client : operates in embedded mode and standalone server
Thrift C++ Client : operated only in embedded mode

3) JDBC Driver :  
For embedded mode, uri is just "jdbc:hive://". 
For standalone server, uri is "jdbc:hive://host:port/dbname" where host and port are determined by where the hive server is run. 
Example :  "jdbc:hive://localhost:10000/default". Currently, the only dbname supported is "default".
JDBC Client Sample Code JDBC Client Sample Code

4) ODBC Driver : 

The Hive ODBC Driver is a software library that implements the Open Database Connectivity (ODBC) API standard for the Hive database management system, enabling ODBC compliant applications to interact seamlessly (ideally) with Hive through a standard interface. This driver will NOT be built as a part of the typical Hive build process and will need to be compiled and built separately ODBC Client

The Metastore :


The metastore is the central repository of Hive metadata. The metastore is divided into two pieces: a service and the backing store for the data. By default, the metastore service runs in the same JVM as the Hive service and contains an embedded Derby database instance backed by the local disk. This is called the embedded metastore configuration




Disadvantage with Embedded metastore is only one Hive Session can be opened against Hive Server, For multiple Hive Session support move metastore to a Relational Database which supports JDO (Derby Server,MYSQL ..) 

MetaStore : Derby Server

wget http://archive.apache.org/dist/db/derby/db-derby-10.4.2.0/db-derby-10.4.2.0-bin.tar.gz

tar -xzf db-derby-10.4.2.0-bin.tar.gz


mv db-derby-10.4.2.0-bin derby


mkdir derby/data


Login as root, add follwing variables in /etc/profile.d/derby.sh



export DERBY_INSTALL=/home/HADOOP/derby
export DERBY_HOME=/home/HADOOP/derby



cd /home/HADOOP/derby/bin/


./startNetworkServer -h 0.0.0.0 &

/scratch/rjuluri/HADOOP/hive/conf

vi hive-site.xml

  hive.test.mode.nosamplelist
 
  if hive is running in test mode, dont sample the above comma seperated list of tables

  hive.metastore.local
  true
  controls whether to connect to remove metastore server or open a new metastore server in Hive Client JVM

  javax.jdo.option.ConnectionURL
jdbc:derby://hostname:port/metastore_db;create=true
  JDBC connect string for a JDBC metastore

  javax.jdo.option.ConnectionDriverName
org.apache.derby.jdbc.ClientDriver
  Driver class name for a JDBC metastore

Add following file 

vi /home/HADOOP/hive/conf/jpox.properties

javax.jdo.PersistenceManagerFactoryClass=org.jpox.PersistenceManagerFactoyImpl
org.jpox.autoCreateSchema=false
org.jpox.validateTables=false
org.jpox.validateColumns=false
org.jpox.validateConstraints=false
org.jpox.storeManagerType=rdbms
org.jpox.autoCreateSchema=true
org.jpox.autoStartMechanismMode=checked
org.jpox.transactionIsolation=read_committed
javax.jdo.option.DetachAllOnCommit=true
javax.jdo.option.NontransactionalRead=true
javax.jdo.option.ConnectionDriverName=org.apache.derby.jdbc.ClientDriver
javax.jdo.option.ConnectionURL=jdbc:derby://hostname:port/metastore_db;create=true
javax.jdo.option.ConnectionUserName=APP
javax.jdo.option.ConnectionPassword=mine

cp /home/HADOOP/derby/lib/derbyclient.jar  /home/HADOOP/hive/lib
cp  /home/HADOOP/derby/lib/derbytools.jar  /home/HADOOP/hive/lib

Hive> show tables;

OK
Time taken: 0.048 seconds

Execute max_cgpa.q script (Previous article Hive - Part I )

[root@slc01mcd hive-0.9.0]# hive -f max_cgpa.q


Hadoop job information for null: number of mappers: 0; number of reducers: 0
2012-07-18 07:01:23,161 null map = 0%,  reduce = 0%
2012-07-18 07:01:26,174 null map = 100%,  reduce = 0%
2012-07-18 07:01:29,186 null map = 100%,  reduce = 100%
Ended Job = job_local_0001
Execution completed successfully
Mapred Local Task Succeeded . Convert the Join into MapJoin
OK


cse     8.6
ece     9.0


Time taken: 10.476 seconds


Hive> show tables;

OK

maxcgpa1

Time taken: 2.769 seconds

Wednesday, July 04, 2012

PIG - HADOOP : Installation , Execution of Sample Program

PIG - HADOOP : Installation , Execution  of Sample Program

Pig raises the level of abstraction for processing large datasets. MapReduce allows you, as the programmer, to specify a map function followed by a reduce function, but working out how to fit your data processing into this pattern, which often requires multiple MapReduce stages, can be a challenge. With Pig, the data structures are much richer, typically being multivalued and nested, and the set of transformations you can apply to the data are much more powerful.


Pig is made up of two pieces:

• The language used to express data flows, called Pig Latin.
• The execution environment to run Pig Latin programs. There are currently two environments: local execution in a single JVM and distributed execution on a Hadoop cluster.


A Pig Latin program is made up of a series of operations, or transformations, that are applied to the input data to produce output. Taken as a whole, the operations describe a data flow, which the Pig execution environment translates into an executable representation and then runs. Under the covers, Pig turns the transformations into a series of MapReduce jobs

Installing and Running Pig

Download latest version of Pig from the following link (Pig Installation).

$ tar xzf pig-0.7.0.tar.gz

set pig environment variables

$ export PIG_INSTALL=/home/user1/pig-0.7.0.tar.gz
$ export PATH=$PATH:$PIG_INSTALL/bin

You also need to set the JAVA_HOME environment variable to point to a suitable Java installation.

Pig has two execution types or modes: 

1) local mode : Pig runs in a single JVM and accesses the local filesystem. This mode is suitable only for small datasets.

$ pig -x local

grunt>

This starts Grunt, the Pig interactive shell

2) MapReduce mode : In MapReduce mode, Pig translates queries into MapReduce jobs and runs them on a Hadoop cluster. The cluster may be a pseudo- or fully distributed cluster.


set the HADOOP_HOME environment variable for finding which Hadoop client to run.

$ pig  or $ pig -x mapreduce , runs pig in MapReduce mode

Running Pig Programs

There are three ways of executing Pig programs, all of which work in both local and MapReduce mode


Script : Pig can run a script file that contains Pig commands. For example, pig
script.pig runs the commands in the local file script.pig
$ pig script.pig

Grunt : Grunt is an interactive shell for running Pig commands.It is also possible to run Pig scripts from within Grunt using run and exec.


Embedded :
You can run Pig programs from Java using the PigServer class, much like you can use JDBC to run SQL programs from Java.

PigPen is an Eclipse plug-in that provides an environment for developing Pig programs.

PigTools and EditorPlugins for pig can be downloaded from PigTools

Example of Pig in Interactive Mode (Grunt)

max_cgpa.pig


-- max_cgpa.pig: Finds the maximum cgpa of a user

records = LOAD 'pigsample.txt'
AS (name:chararray, spl:chararray, cgpa:float);
filtered_records = FILTER records BY cgpa > 0 AND cgpa < 10;
grouped_records = GROUP filtered_records BY spl;
max_cgpa = FOREACH grouped_records GENERATE group, MAX(filtered_records.cgpa);
STORE max_cgpa INTO 'output/cgpa_out';

Above pig script finds the maximum cgpa of a specialization.

pigsample.txt  ( Input to the pig )

raghu     ece     9
kumar    cse      8.5
biju       ece      8
mukul    cse      8.6
ashish   ece      7.0
subha    cse      8.3
ramu     ece     -8.3
rahul     cse      11.4
budania ece      5.4

first column represents name , second column specialization and third column is cgpa, by default each column is separated by tab space.

$ pig max_cgpa.pig

Output : 

(cse,8.6F)
(ece,9.0F)

Analysis : 

Statement : 1
records = LOAD 'pigsample.txt'AS (name:chararray, spl:chararray, cgpa:float);

Load input file in to memory from the file system (HDFS or local or Amazon S3). name:chararray notation describes the field’s
name and type; chararray is like a Java string, and an float is like a Java float.

grunt> DUMP records;

(raghu,ece,9.0F)
(kumar,cse,8.5F)
(biju,ece,8.0F)
(mukul,cse,8.6F)
(ashish,ece,7.0F)
(subha,cse,8.3F)
(ramu,ece,-8.3F)
(rahul,cse,11.4F)
(budania,ece,5.4F)

Input is converted in to a tuple , and each column is separated by ,

grunt> DESCRIBE records;
records: {name: chararray,spl: chararray,cgpa: float}

Statement : 2
filtered_records = FILTER records BY cgpa > 0 AND cgpa < 10;

grunt> DUMP filtered_records;

filter all the records whose cgpa <0 (negative) and >10 

(raghu,ece,9.0F)
(kumar,cse,8.5F)
(biju,ece,8.0F)
(mukul,cse,8.6F)
(ashish,ece,7.0F)
(subha,cse,8.3F)
(budania,ece,5.4F)

grunt> DESCRIBE filtered_records;
filtered_records: {name: chararray,spl: chararray,cgpa: float}

Statement : 3

The third statement uses the GROUP function to group the records relation by the specialization field.

grouped_records = GROUP filtered_records BY spl;

grunt> DUMP  grouped_records ;

(cse,{(kumar,cse,8.5F),(mukul,cse,8.6F),(subha,cse,8.3F)})
(ece,{(raghu,ece,9.0F),(biju,ece,8.0F),(ashish,ece,7.0F),(budania,ece,5.4F)})

grunt> DESCRIBE  grouped_records;
grouped_records: {group: chararray,filtered_records: {name: chararray,spl: chararray,cgpa: float}}

We now have two rows, or tuples, one for each specialization in the input data. The first field in each tuple is the field being grouped by (the specialization), and the second field is a bag of tuples
for that  specialization. A bag is just an unordered collection of tuples, which in Pig Latin is represented using curly braces.

By grouping the data in this way, we have created a row per  specialization , so now all that remains is to find the maximum cgpa for the tuples in each bag.

Statement : 4


max_cgpa = FOREACH grouped_records GENERATE group,
MAX(filtered_records.cgpa);

FOREACH processes every row to generate a derived set of rows, using a GENERATE clause to define the fields in each derived row. In this example, the first field is group, which is just the specialization. The second field is a little more complex.

The filtered_records.cgpa reference is to the cgpa field of the
filtered_records bag in the grouped_records relation. MAX is a built-in function for calculating the maximum value of fields in a bag. In this case, it calculates the maximum cgpa for the fields in each filtered_records bag.

grunt> DUMP    max_cgpa  ;

(cse,8.6F)
(ece,9.0F)

grunt> DESCRIBE    max_cgpa  ;

max_cgpa : {group: chararray,float}

Statement : 5

STORE max_cgpa INTO 'output/cgpa_out'

This command redirects the output of the script to a file (Local or HDFS) instead of printing the output on the console .

we’ve successfully calculated the maximum cgpa for each specialization.

With the ILLUSTRATE operator, Pig provides a tool for generating a reasonably complete and concise sample dataset.


--------------------------------------------------------------------
| records     | name: bytearray | spl: bytearray | cgpa: bytearray | 
--------------------------------------------------------------------
|             | kumar           | cse            | 8.5             | 
|             | mukul           | cse            | 8.6             | 
|             | ramu            | ece            | -8.3            | 
--------------------------------------------------------------------
----------------------------------------------------------------
| records     | name: chararray | spl: chararray | cgpa: float | 
----------------------------------------------------------------
|             | kumar           | cse            | 8.5         | 
|             | mukul           | cse            | 8.6         | 
|             | ramu            | ece            | -8.3        | 
----------------------------------------------------------------
-------------------------------------------------------------------------
| filtered_records     | name: chararray | spl: chararray | cgpa: float | 
-------------------------------------------------------------------------
|                      | kumar           | cse            | 8.5         | 
|                      | mukul           | cse            | 8.6         | 
-------------------------------------------------------------------------
----------------------------------------------------------------------------------------------------------------
| grouped_records     | group: chararray | filtered_records: bag({name: chararray,spl: chararray,cgpa: float}) | 
----------------------------------------------------------------------------------------------------------------
|                     | cse              | {(kumar, cse, 8.5), (mukul, cse, 8.6)}                              | 
----------------------------------------------------------------------------------------------------------------
-------------------------------------------
|  max_cgpa   | group: chararray | float | 
-------------------------------------------
|              | cse              | 8.6   |


EXPLAIN max_cgpa  

Use the above command to see the logical and physical plans created by Pig.

Tuesday, February 21, 2012

How to Surrender ICICI prudential Life Insurance(ULIP)

How to Surrender ICICI prudential Life Insurance       (ULIP)


If you Opted for ICICI Life Insurance (ULIP) for Tax Saving 80C, there is a lock in period of 3 years , so every year you need to pay the premium until 3 years , after completion of 3 years you are eligible for full Withdrawal or Partial withdrawal, for full Withdrawal ICICI charges 2% on the outstanding amount.


Fill up the Payout request for Surrender / Partial Withdrawal form from the ICICI Prulife site ICICI Prulife Surrender Form




Following are the Mandatory documents for withdrawal


1) Self attested Photo-ID proof (Driving License/Passport..)
2) Copy of signed cancelled cheque ( amount will be transferred to this account)
3) Original Policy Certificate
4) ICICI Prulife Surrender Form


above 4 documents are mandatory , and this needs to be submitted in any ICICI  Prulife branch.



If the application for re-instatement and surrender is received on the same day, first the policy will be re-instated and then the surrender will be processed on the next working day  and the NAV of the date of processing will be applicable.



Popular Posts