We provide IT Staff Augmentation Services!

Hadoop/spark Developer Resume

2.00/5 (Submit Your Rating)

Fremont, CA

SUMMARY

  • 7 years of extensive Professional IT experience, including 5 years of Hadoop/Bigdata experience, capable of processing large sets of structured, semi structured and unstructured data and supporting systems application architecture.
  • Deep understanding of Hadoop Architecture of versions 1x,2x and various components such as HDFS, YARN and MapReduce concepts along with Hive, Pig, Kafka, Sqoop, Oozie, Zookeeper, Map Reduce framework and NoSQL databases like HBase.
  • Worked on development projects that were based on Data Virtualization concepts and creating web services in Denodo.
  • Experience in Scala’s FP, Case Classes, Traits and leveraged Scala to code Spark applications.
  • Expertise in writing Hive and Pig scripts and UDFs to perform data analysis on large data sets.
  • Hands on experience in installation, configuration, management and deployment of Big Data solutions and the underlying infrastructure of Hadoop Cluster using Cloudera, Hortonworks distributions and MapR.
  • Managed and Scheduled jobs on Hadoop cluster using Apache Oozie.
  • Partitioned and Bucketed data sets in Apache Hive to improve performance.
  • Hands on experience in configuring and working with Flume to load the data from multiple sources directly into HDFS and transferred large datasets between Hadoop and RDBMS by implementing SQOOP.
  • Perform transformations like event joins, filter bot traffic and some pre - aggregations using Pig.
  • Develop MapReduce jobs to convert data files into Parquet file format.
  • Execute Hive queries on Parquet tables stored in Hive to perform data analysis to meet the business requirements.
  • Develop business specific Custom UDF's in Hive, Pig.
  • Integrated Presto with MySQL and Hive following the steps outlined in theMySQLconnector andHivedocumentation respectively.
  • Experienced in writing queries using Presto for group by and sort by on the tables.
  • Used Presto with Hive meta store to query the tables in hive metastore.
  • Experience in performing SQL and hive operations using Spark SQL.
  • Responsible for spark streaming using Scala API and process data stored in NOSQL data bases
  • Developed Spark SQL scripts and involved in converting hive UDF’s to Spark SQL UDF’s
  • Develop a data pipeline using Kafka to store data into HDFS.
  • Responsible for batch processing by creating hive context using Spark SQL and push data set into Datawarehouse(DWS) for further processing uses of tableau (BI team) as well as data science team.
  • Good knowledge on TalendDQ & Data profiling.Responsible for batch processing and real time processing in HDFS and NOSQL Databases.
  • Trouble shooting, debugging & altering Talendparticular issues, while maintaining the health and performance of the ETL environment.
  • Extensively created mappings in Talendusing t-Map, t-Join, t-Replicate, t-Parallelize, t-Die, t-Aggregate Row, t-Warn, t-Log Catcher, t-Filter, t-Global map etc.
  • Experienced in working with various kinds of data sources such asHortonworksTeradata and Oracle.
  • Utilized Apache Hadoopenvironment by Hortonworks.
  • Experience with cluster management technologies such as YARN, Mesos.
  • Experienced in installation, configuration, support and monitoring of Hadoop clusters using Cloudera distributions, Hortonworks HDP, MapR and AWS.
  • Experience in managing Hadoop clusters using Cloudera Manager Tool and Hue.
  • Created Kafka Topics and distributed to different consumer applications.
  • Developed applications using Java, RDBMS and UNIX Shell scripting.
  • Hands-on experience with visualization tools such as MS Excel, MS Visio, and Tableau.
  • Experience in Basic Hadoop administration such as replication, node removal.
  • Great knowledge in documenting the Software Requirements Specifications including Functional Requirements, Data Requirements and Performance Requirements.
  • Highly organized with the ability to manage multiple projects and meet deadlines.

TECHNICAL SKILLS

Hadoop/Big Data: HDFS, MapReduce, Pig, Hive, Sqoop, Oozie, Apache Spark, Flume, Kafka, Scala, Impala, Presto

Distributions: MapR, Cloudera, AWS, Hortonworks

No SQL Databases: HBase, Cassandra, MongoDB

Development/Build Tools: Eclipse 4.4/4.3/3.8/3.7/3.5/3.2 Net Beans, Edit Plus 2, Ant, Maven, Gradle, IntelliJ, JUNIT and log4J., Spring MVC, Hibernate

Development Methodology: Agile, Unified Modeling Language (UML), Design Patterns (Core Java and J2EE), RationallyUnifiedProcess, WaterFall, Iterative.

Programming languages: SQL, Linux shell scripts, C, Java

Databases: MySQL, DB2, ODBC

Database Languages: MySQL, PL/SQL, Oracle

RDBMS: Teradata, Oracle … MS SQL Server, MySQL and DB2

ETL Tools: MS Office suite, RAW, Tableau, Talend

Operating Systems: UNIX, LINUX, Mac OS and Windows Variants

Web Services: WebSphere, WebLogic, JBoss and Tomcat

File Formats: XML, Text, Sequence, RC, JSON, ORC, AVRO, and Parquet etc.

PROFESSIONAL EXPERIENCE

Confidential, Fremont CA

Hadoop/Spark Developer

Responsibilities:

  • Good knowledge and experience with Hadoop stack - internals, Hive, Pig and Map Reduce
  • Hands on experience in loading source data like Web Logs using Kafka pipelining to HDFS.
  • Experienced on loading and transforming of large sets of data from Cassandra source through Kafka and placed in HDFS for further processing.
  • Created Hive Tables, loaded transactional data from RDBMS using Kafka
  • Collected and aggregated large amounts of web log data from different sources such as webservers, mobile and network devices using Apache Kafka and stored the data into HDFS for analysis.
  • Developed Spark SQL scripts and involved in converting hive UDF’s to Spark SQL UDF’s
  • Responsible for batch processing and real time processing in HDFS and NOSQL Databases.
  • Load the data into Spark RDD and do in memory data Computation to generate the Output response.
  • Developed Spark scripts by using Scala shell commands as per the requirement
  • Used Spark API over ClouderaHadoopYARN to perform analytics on data in Hive
  • Created Partitioned and Bucketed Hive tables in Parquet File Formats with Snappy compression and then loaded data into Parquet hive tables from Avro hive tables
  • Experienced in performance tuning of Spark Applications for setting right Batch Interval time, correct level of Parallelism and memory tuning
  • Optimizing of existing algorithms inHadoopusing Spark Context, Spark-SQL, Data Frames and Pair RDD's
  • Load the data into Spark RDD and performed in-memory data computation to generate the output response
  • Load Data into Hbase using Bulk Load and Non-bulk load.
  • Responsible for batch processing by creating hive context using Spark SQL and push data set into Data warehouse(DWS) for further processing uses of tableau (BI team) as well as data science team.
  • Responsible for spark streaming using Scala API and process data stored in NOSQL data bases
  • Experienced with performing analytics on Time Series data using HBase
  • Implemented HBase co-processors, Observers to work as event based analysis
  • Extensive Working knowledge of partitioned table, UDFs, performance tuning, compression-related properties, thrift server in Hive.
  • Performed operation using Partitioning pattern in Map Reduce to move records into different categories
  • Experienced in working with Amazon Web Services (AWS) using EC2 for computing and S3 as storage mechanism.
  • Implemented Hive Generic UDF's to implement business logic
  • Experienced with accessing Hive tables to perform analytics from java applications using JDBC.
  • Installed and configured Hive and written Hive QL scripts. Experienced with multiple file formats in HIVE, like Sequence file format, ORC file format etc.,
  • Implemented business logic by writing Pig UDF's in Java and used various UDFs from Piggybanks and other sources
  • Extensive experience in writing UNIX shell scripts and automation of the ETL processes using UNIX shell scripting.
  • Responsible for coding Java Batch, Restful Service, Map Reduce program, Hive query’s, testing, debugging, Peer code review, troubleshooting and maintain status report.
  • Worked with Developers in ensuring technical design gets translated into ETL Graphs, review the code in making sure it gets developed as per DDE Design/Development guidelines.
  • Co-ordinate with Quality Services in making sure they understand the requirement clearly in order for their validations.

Environment: Casandra, Map jobs, Spark 1.6.3, Spark SQL, Pig Scripts, Pig UDF’s, Oozie, HIVE, AVRO, Hive, Scala, Kafka, Map Reduce, Restful Service, Java, Eclipse, AWS, Python, Unix, Oracle DB.

Confidential, New Jersey NJ

HADOOP DEVELOPER

Responsibilities:

  • Sqoop jobs, PIG and Hive scripts were created for data ingestion from relational databases to compare with historical data
  • Responsible for building scalable distributed data solutions using Hadoop
  • Transformed incoming data with Hive & Pig to make data available to internal users
  • Performed extensive Data Mining applications using HIVE
  • Worked on Sequence files, RC files, Map side joins, bucketing, partitioning for Hive performance enhancement and storage improvement
  • Process the data and push the valid records to HDFS.
  • Import data from MySQL to HDFS using SQOOP.
  • Tune the MapReduce, PIG and Hive jobs to increase the performance and decrease the execution time of the jobs.
  • Compress the files downloaded from the servers before storing them in the cluster to save cluster resources.
  • Write Corejava programs to convert the JSON files to CSV or TSV files for further processing.
  • Optimize already developed long running MapReduce and Pig job for better performance and accurate results.
  • Create Hive databases and tables over the HDFS data and write HiveQL queries on the tables.
  • ScheduleHadoopand UNIX jobs using OOZIE.
  • Work with NoSQL databases like HBase.
  • Write Pig and HiveUDFs for processing and analyzing log files.
  • Developing Scripts and Batch Job to schedule variousHadoopProgram.
  • Visualize the complicated data analysis on the dashboards as per the business requirements.
  • Integrated Hive, PIG and Mapreduce jobs with elastic search to publish the metrics to the dashboards.
  • Utilized the most usedTalendComponents such as tMap, tFilterRow, tAggregateRow, tFileExist, tFileCopy, tFileList, tDie etc.
  • Also, Utilized Big Data components such as tSqoopExport, tSqoopImport, tHDFSInput, tHDFSOutput, tHiveLoad, tHiveInput, tPigLoad, tPigFilterRow, tPigFilterColumn, tPigStoreResult,, tHbaseInput, tHbaseOutput along with executing the jobs in Debug mode and also utilizing the tlogrow component to view the sample output.
  • Submittedtalendjobs for scheduling usingTalendscheduler which is available in the Admin Console.
  • Deployedtalendjobs on various environments including dev, test and production environments.
  • Involved in analysis, design testing phases and responsible for documenting the technical specifications.

Environment: Hadoop2x, YARN, HDFS, MapReduce, PIG, HIVE, HBASE, Shell Scripting, java, Oozie,TALEND, LINUX.

Confidential, New York NY

HADOOP DEVELOPER

Responsibilities:

  • Collected log data and staging data using Apache Flume and stored in HDFS for analysis.
  • Implemented helper classes that access HBase directly from java using Java API to perform CRUD operations.
  • Handled different time series data using HBase to perform store data and perform analytics based on time to improve queries retrieval time.
  • Developed MapReduce programs to parse the raw data and store the refined data in tables.
  • Performed debugging and fine tuning in Hive & Pig for improving performance.
  • Used Oozie operational services for batch processing and scheduling workflows dynamically.
  • Analyzed the web log data using the HiveQL to extract number of unique visitors per day.
  • Exported the analysed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
  • Performed Map side joins on data in Hive to explore business insights.
  • Involved in forecast based on the present results and insights derived from data analysis.
  • Integrated Map Reduce with HBase to import bulk amount of data into HBase using Map Reduce Programs.
  • Participated in team discussions to develop useful insights from big data processing results.
  • Suggested trends to the higher management based on social media data.

Environment: HDFS, MapReduce, Hive, HBase, Pig, Java, Git, Maven, Talend, Putty, REST, CentOS 6.3

Confidential, San Francisco, CA

Hadoop Developer

Responsibilities:

  • Installed and configured Hadoop Map Reduce, HDFS, developed multiple Map Reduce jobs in java for data cleaning and preprocessing
  • Supported in setting up updating configurations for implementing scripts with Pig and Sqoop
  • Migrated existing SQL queries to HiveQL queries to move to big data analytical platform
  • Integrated Cassandra file system to Hadoop using Map Reduce to perform analytics on Cassandra data
  • Wrote test cases in Junit for unit testing of classes
  • Used Hibernate ORM framework with Spring framework for data persistence
  • Responsible to manage data coming from different sources
  • Supported the clusters that were running on Map Reduce Programs
  • Involved in loading data from UNIX file system to HDFS
  • Involved in templates and screens in HTML and JavaScript
  • Load and transform large sets datainto HDFS using Hadoop fs commands
  • Designed the logical and physical data modeling wrote DML scripts for Oracle 9i database
  • Experience with installing cluster, commissioning as well as decommissioning of different types of nodes like Name and Data nodes

Confidential

Java Developer

Responsibilities:

  • Involved in analysis and gathering requirements and user specifications from business analyst.
  • Involved in creating use case, class, sequence, package dependency diagrams using UML.
  • Involved in Database Design by creating Data Flow Diagram (Process Model) and ER Diagram (Data Model).
  • Used JavaScript for certain form validations, submissions and other client side operations.
  • Developed application business components and configured beans using Spring IOC.
  • Generated POJO classes and Hibernate mapping files using Reverse Engineering.
  • Developed DAO classes using Hibernate Template from spring with Hibernate API.
  • Designed and Implemented MVC architecture using Spring MVC.
  • Created Stateless Session Beans to communicate with the client.
  • Created Connection Pools and Data Sources.
  • Implemented and supported the project through development, Unit testing phase into production environment.
  • Designing the database and coding of SQL, PL/SQL, Triggers and Views using IBM DB2.

Environment: Java 5.0, JavaScript DB2, Windows XP, Eclipse IDE, Web Logic Server, SQL, PL/SQL.

We'd love your feedback!