Hadoop Developer And Big Data Analyst Resume
Bellevue, WA
SUMMARY:
- 8+ Years of IT experience in Development& Enhancement of Data warehousing and Hadoop ECO systems, as a developer on Mainframes Programming, Configuration Engineer on Endevor Configuration Management, Unix Prod Support in AIX &Solaris,
- Had 3+ plus years of experience as Hadoop Developer and Big data analyst. Expertise in writing Hadoop Jobs for analyzing data using HDFS, Hive, HBase, Pig, Sqoop, IMPALA, YARN, Spark, Scala, AWS Storage and had good knowledge on Oozie, MONGODB, HBASE.
- Had Comprehensive Experience in Development & Enhancement of Data warehouses and have worked extensively on Informatica 9.5.1, UNIX,Tidal and Teradata 14.
- Had 3 year and 3months of IT experience of complete lifecycle implementations from design and migration of the elements as a Configuration Engineer through ENDEVOR TOOL and Unix Prod Support.
- Had In - depth Expertise in Telecom Domain.
- Good knowledge of Hadoop Architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and MapReduce concepts.
- Working experience on designing and implementing complete end-to-end Hadoop Infrastructure including PIG, HIVE, Sqoop, Flume/ Kafka.
- Performed Sales to Billing Analysis on Video Applications and Churn Analysis using Data on Hadoop Cluster.
- Good knowledge with Amazon Web Services (AWS) tools such as Amazon S3
- Good Knowledge in AWS concepts like EMR and EC2, AWS redshift, Dynamo DB, AWS S3 storage, AWS SNS which provides fast and efficient processing of Big Data.
- Had exposure and understanding skills on Apache NIFI, Falcon and Ranger.
- Had experience leading team for offshore and onsite coordination.
- Good understanding in installation, configuration of Apache Hadoop and Cloudera CDH clusters.
- Had exposure and used MAP REDUCE Programs
- Excellent communication, demonstrated interpersonal and leadership skills
- Good Knowledge of the Configuration Management tool -Endeavor,
- Good knowledge on Solaris 10, IBM AIX 5.3L, Backup Administration
- Good Knowledge on DB2-DBA and DB2-UDB admin concepts.
TECHNICAL SKILLS:
BIG DATA: Apache Hadoop 2.6.5,CloudEra Enterprise 5.X able to handle Hadoop on Horton Works
NO SQL databases: HBASE 0.94.5, MONGO DB, Cassendra
Hadoop Ecosystems: HIVE 1.2, PIG 0.11.0, Sqoop 1.4.3, Flume 1.4.0 and KAFKA
Operating Systems: Z/OS, Solaris 10, IBM AIX 5.3L, Ubuntu 13.X
Languages: SCALA, COBOL,SQL, R language and Core java and Python
Databases: DB2-DBA, DB2-UDB, Teradata 14, HIVE
Configuration TOOLS: ENDEVOR
ETL TOOLS: Informatica 9.5.1
Applications: MS office
Development Tools: SSH Secure Shell, Teradata SQL Assistance
Scheduling Tools: Tidal, ESP, CRON TAB on UNIX
PROFESSIONAL EXPERIENCE:
Confidential, Bellevue, WA
Hadoop Developer and Big data analystResponsibilities:
- Develop New Spark Sql ETL logics in Big Data for the migration and availability of the Facts andDimensions used for the Analytics
- Develop of Spark Sql application, Big Data Migration from Teradata to Hadoop and reduce Memory utilization in Teradata analytics.
- Requirement Gathering and Leading Team for the development of the Big Data environment and Spark ETL logics migrations.
- Involve in requirement gathering from the Business Analysts, and participate in discussions with users, functional analysts for the Business logics implementation.
- Responsible for end to end design on Spark Sql, Development to meet the requirements.
- Advice the business on best practices in the Spark Sql while making sure the solution meet the business needs.
- Lead and Coordinate Developers, Testing and technical teams in offshore support on daily basis to discuss Challenges and outstanding issues.
- Involve in preparation, distribution and collaboration of client specific quality documentation on developments for Big Data and Spark along with regular monitoring on reflecting the modifications or enhancements done in Confidential Schedulers.
- Migrate the Data from Teradata to Hadoop and data preparation using HIVE Tables.
- Create Partitioned and bucketing tables on HIVE.Mainly worked on Hive QL to categorize data of different Subject areas for Marketing, Shipping, and Selling.
- Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.
- Accessing the hive tables using Spark Hive context (Spark sql) and used scala for interactive operations.
- Develop the Spark Sql logics which mimics the Teradata ETL logics and point the output Delta back to Newly Created Hive Tables and as well the existing TERADATA Dimensions, Facts, and Aggregated Tables.
- Make sure Data is matched with TERADATA and SPARK Sql logics.
- Creating Views on Top of the HIVE tables and give it to customers for the analytics.
- Good Knowledge on KYLIN which is used as MOLAP engine for analytical usage.
- Had worked on different input formats like Clickstream, XML, Jason,text files.
- Created External tables with Sequence, Avro, and Parquet Format in Hive.
Environment: Hive, SPARKSql, Hadoop, HDFS, Teradata 14, SQL Assistant, MultiLoad, FastLoad, BTEQ, FastExport,.
Confidential, Irving, TX
Analyst -Sr. Bigdata Hadoop Developer
Responsibilities:
- Identify the Customer Usage, Behavior, Feedback datasets which would reside on different systems and unfolding insights in to Customer Usage, behavior and Feedback analysis who are likely to Churn.
- Analyzed data which need to be loaded into hadoop and contacted with respective source teams to get the table information and connection details.
- Migrating the data of high volumes from Oracle, MySQL, Teradata in to HDFS using Sqoop, Informatica ETL and importing various formats of flat files in to HDFS.
- Creating Partitioned tables HIVE.Mainly worked on Hive QL to categorize data of different claims, Implemented Partitioning, Dynamic Partitions, Buckets in HIVE.
- Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.
- Working on projecting, involving and migration of data from different sources, Teradata to HDFS Data Lake and creating reports by performing transformations on the data put in the Hadoop data lake.
- Cleansing the Data and developinga strategy for Full load and incremental load using Sqoop.
- Analyze data in Pig Latin, Hive and usedMap Reduce programs in Java.
- Had good exposure and knowledge on YARN framework, IMPALA
- Had worked on SPARK SQL and had good exposure on SPARK components
- Had knowledge on SCALA and worked on the Transformations, Actions
- Had good experience on accessing the hive tables using Spark Hive context (Spark sql) and used scala for interative operations.
- Had exposure and worked with streaming technologies like Spark Streaming with Kafka, SparkSQL.
- Used RDD's to perform transformation on datasets as well as to perform actions like count, reduce, first.
- Import tables from different databases to HDFS and HBase using Sqoop.
- Had exposure on workflows for Pig and Hive jobs in Oozie.
- Had worked on different input formats like Clickstream, XML, Jason,text files, Sequence files, Avro data files using SerDe's in Hive.
- Had good exposure on AWS Cloud and using the instances of S3 storage for retrieving the archival data.
- Good knowledge and exposure on Dynamo DB, AWS EMR, EC2, AWS S3 storageetc
- Gained knowledge in integrating the hive warehouse with Hbase, AWS.
- Constantly worked on tuning the performance of the queries in Hive and dataflow in Pig, making the queries work even more powerfully in processing and retrieving the data
- Along with the Infrastructure team, involved in design and developed Kafka and Streaming data pipeline
- Had knowledge on R language and have used it on Proof of concepts in Datascience statistical models.
- Documented ETL best practices to be implemented with Hadoop
- Monitoring and Debugging Hadoop jobs/Applications running in production.
- As a team member monitored and supported on Hadoop Cloudera upgrade from CDH4.X to CDH5. x.
- Had good knowledge on the Integration of Hadoop, Hive with AWS, HBASE etc..
- Develop HIVEand PIG Jobs and Worked with Data-scientists on Sentimental Analysis to identify the locations where percentage of CHURN is high and identify customers performing rotational Churns.
- Had good exposure on NOSQL Databases like MONGODB, HBASEand gaining knowledge on Cassendra.
- Preparing the Design, Approach and Solution documents and Follow the SDLC cycle for production implementations.
Environment: Oozie, Sqoop, Hive, Pig, Python, SCALA, SPARK, HBASE, Core java,Hadoop, HDFS, MONGODB, Teradata 14, SQL Assistant, MultiLoad, FastLoad, BTEQ, FastExport, AWS and NIFI data flow language.
Confidential
Analyst- sr. Data warehouse Developer
Responsibilities:
- Extracting, transforming and loading data from Various sources on Daily, weekly and monthly basis into the database with the help of batch jobs to ensure the data loaded is 100 % fine.
- Supported Developed mappings& updating the existing mappings as per the requirement.
- Analysis of existing code and testing of the programs as per the requirement.
- Monitoring the Incident Management queue &fixing the issues with appropriate resolution steps and maintaining proper documentation for further reference.
- Comparing the record counts with the source file counts by using Informatica monitor.
- Involved in Datamart data preparation and Data cleansing before the Data reaching into DataWarehouse.
- Involved creating tickets, creating Change Activity documents etc., when the New Changes in code has to promote to production.
- Used most of the transformations such as the Source Qualifier, Expression, Aggregator, Connected & unconnected lookups, Filter, Router, Sequence Generator, Sorter, Joiner, and Update Strategy.
- Imported data from various Sources transformed and loaded into Data Warehouse Targets using Informatica.
- Extensively worked on data extraction, transformation and data loading from source to target system using Teradata Utilities BTEQ and Fastload, MLOAD according to Business needs.
- Documentation (LLD) to describe program development, logic, coding, testing, changes and corrections.
- Interacting with business users to let them know about the availability of data and imbalanced data.
- Monitor the Daily, Weekly and monthly production loads in Tidal Scheduler.
Environment: Informatica, Teradata 14, SQL Assistant, MultiLoad, FastLoad, BTEQ, FastExport.
Confidential
Member Technical Staff
Responsibilities:
- Attending e - mail queries from the users,alerts on CMAT Confidential Proprietary web tool.
- Executing Code Cut to Dev Environment, Scheduled and Immediate distributions rollthe code intestEnvironments and Upgrading Productionenvironments.
- Working on Differentrequests.Archiving the Prod Code Quarterly and Regular release.
- Maintaining 24x7 CM control desk with Endevor and ESP scheduling tool
- Attending client meetings, turnover meetings and status calls and Coordinating the configuration management activities with theoffshore/Onshoreteam.
- Reviewing and providing feedback to the work done by offshore.
- Interacting with users and help them fix any Endevor or Configuration issues.
Environment: Endevor, ESPscheduler, COBOL, DB2, CMAT inbuit Confidential configuration GUI Tool
Confidential
Software Engineer
Responsibilities:
- Working on Create and publish the Master Element List (MEL). Notify Release Team for list availability
- Developed and Enhanced the code based on the requirement.
- Performed the Unit test and make sure the requirement was met.
- Work involved and to Ensure Development has Retrieved, w/Sign Out, all the elements on the Master Element List with the correct CCID.
- Doing the Pre-Implementation activities and involving in the Implementation activities and Work on Delete requests on customer request.
- Attending the Release Kick off meetings, working along with the release schedule and following up the development teams for proper sign into the Release Paths.
Environment: Endevor, ESPscheduler, COBOL, DB2.
