Hadoop Developer Resume
Southborough, MA
SUMMARY
- Over 7+ years of IT experience in full System Development Life Cycle using WATERFALL and AGILE methodologies.
- Over 3+ years of experience working in Big Data Ecosystem related technologies with full project development, implementation and deployment on Linux/Windows/Unix
- Strong experience on Hadoop working environment includes Map Reduce, HDFS, HIVE, PIG, HBase, Zookeeper, Sqoop, Oozie, Spark, Flume and Avro for data storage and analysis
- Experience in shell and python scripting languages
- Very strong experience in processing, analyzing large sets of structured, semi - structured and unstructured data and supporting systems application architecture
- Extensive experience in writing Hadoop jobs for data analysis as per teh business requirements using Hive and Pig
- Expertise in creating Hive Internal/External Tables/Views using shared Meta store
- Developed custom UDFs in Pig and Hive to extend their core functionality
- Extensive experience in handling semi structured/unstructured data using Map Reduce Programs.
- Experience on building real time applications using SPARK streaming.
- Experience in handling different file formats like text files, Sequence files, Avro data files, mahout xml files using Map Reduce programs.
- Hands on experience in transferring incoming data from various application servers into HDFS, Hive, HBase using Apache Flume
- Worked extensively on SQOOP to import and export data from RDBMS to HDFS and vice-versa
- Performed Data Ingestion from multiple disparate sources and systems using Kafka
- Worked on Oozie to manage and schedule teh jobs on Hadoop cluster
- Experience in managing and reviewing Hadoop log files
- Worked with NoSQL database HBase to create tables and store data
- Experience in installation, configuration, support and monitoring of Hadoop clusters using Apache, Cloudera distributions and AWS
- Experience in setting up Hive, Pig, HBase, and SQOOP on Ubuntu Operating system
- Experience in managing Hadoop clusters using Ambari
- Proficient in using data visualization tools like Tableau and MS Excel
- Strong experience in writing UNIX shell scripts
- Working in different projects provided exposure and good understanding of different phases in SDLC
- Experience in developing component design using UML, Use case, Class, Sequence, Deployment and Component diagrams for teh given requirements
- Experience of working on projects in Agile methodology
- Ability to has clear understanding of teh business requirements and use teh technical expertise in arriving at teh best possible solution
TECHNICAL SKILLS
Hadoop/Big Data: Hadoop 1x/2x(Yarn), HDFS, Map Reduce, Hive, Pig, Sqoop, Flume, Kafka, Spark, Storm, Zookeeper, Oozie, Ambari, Tez
Development Tools: Eclipse, IBM DB2 Command Editor, TOAD, SQL Developer, VM Ware
Programming/Scripting Languages: Python, Core Java, Python, SQL, Pig Latin, Hive QL
Databases: Oracle 11g,10g,9i, MySQL, SQL Server 2005,2008, PostgreSQL& DB2
NoSQL Databases: HBase, Cassandra, Mongo DB
Visualization: Tableau, Raw and MS Excel
Version Control Tools: Sub Version (SVN), Concurrent Versions System (CVS) and IBM Rational ClearCase
Methodologies: Agile/ Scrum, Waterfall
Operating Systems: Windows, Unix, Linux and Solaris
PROFESSIONAL EXPERIENCE
Confidential - Southborough, MA
Hadoop Developer
Responsibilities:
- Designing technical architecture and developed various Big Data workflows using custom Spark, Hive and SQOOP.
- Supporting and building teh Data Science team projects on to Hadoop
- Advising teh industrial guidelines on teh Hadoop applications setup
- Benchmarking various options that are available in Hive, Spark and Impala.
- Optimizing existing applications performance.
- Analyzing new Hadoop ingest tools like TALEND, RCG framework
- Built re-usable Hive UDF libraries for business requirements which enabled various business analysts to use these UDF’s in Hive querying.
- Built SPARK pipelines and work flows
- Building SPLR Dashboards for audit data by build indexes on teh Oracle audit data.
- Created custom python/shell scripts to import data via SQOOP from Oracle databases.
- Created big data workflows to ingest teh data from various sources to Hadoop using OOZIE and these workflows comprises of heterogeneous jobs like Hive, SQOOP and Python Script.
- Created various scripts to import teh data various internal sources using Curl commands.
- Converted various SAS scripts into their equivalent Hive QLs.
- Providing technical solutions/assistance to all development projects
- Advising various teams on teh impact of new enhancements/products
Environment: Hortonworks, Hive, SQOOP, FLUME, Apache SPARK, HBASE, JDK 1.7, SOLR, Red hat Linux, Python, Shell scripting, Oracle
Confidential - Cleveland, OH
Hadoop developer
Responsibilities:
- Designing technical architecture and developed various Big Data workflows using custom MapReduce, Hive and SQOOP.
- Built re-usable Hive UDF libraries for business requirements which enabled various business analysts to use these UDF’s in Hive querying.
- Used FLUME to dump teh transaction logs into HDFS.
- Built SPARK pipelines and work flows
- Created custom python/shell scripts to import data via SQOOP from various SQL databases such as Teradata, SQL Server, and Oracle.
- Created big data workflows to ingest teh data from various sources to Hadoop using OOZIE and these workflows comprises of heterogeneous jobs like Hive, SQOOP and Python Script.
- Created various scripts to import teh data various internal sources using Curl commands.
- Created Custom FTP job to import teh mainframe file directly to Hadoop.
- Created Custom Hive SERDE to map teh contents of mainframe into Hive Structure which involves complex mapping of OCCURS, DEPENDING, REDFINES etc. clauses.
- Converted various SAS scripts into their equivalent Hive QLs.
- Build a month-end processing ETL scripts for RREAP system using hive and hive streaming.
- Assigned teh tasks of resolving defects found in testing teh new application and existing applications.
- Analyzing teh requirements, designing and developing solutions.
- Managing Project team in achieving teh project goals including resource allocation, resolving technical issues and mentoring teh resources.
- Providing technical solutions/assistance to all development projects
- Advising various teams on teh impact of new enhancements/products
Environment: Hive, SQOOP, FLUME, Apache SPARK, HBASE, JDK 1.6, Maven, Red hat Linux, Python, Shell scripting, Teradata, SQL Server, Oracle, MySQL
Confidential
Hadoop developer
Responsibilities:
- Responsible to manage data coming from different sources and involved in HDFS maintenance and loading of structured and unstructured data
- Designing technical architecture and developed various Big Data workflows using custom MapReduce, Pig, Hive and SQOOP.
- Deployed on premise cluster and tuned teh cluster for optimal performance for job execution needs and processes large data sets.
- Built re-usable Hive UDF libraries for business requirements which enabled various business analysts to use these UDF’s in Hive querying.
- Used FLUME to dump teh application server logs into HDFS.
- Teh logs that are stored on HDFS are analyzed and teh cleaned data is imported into Hive warehouse which enabled end business analysts to write Hive queries.
- Configured various big data workflows ton run on teh top of Hadoop using OOZIE and these workflows comprises of heterogeneous jobs like Pig, Hive, SQOOP and MapReduce.
- Experience in working with NoSQL database HBASE in getting real time data analytics.
- Developed suit of Unit Test Cases for Mapper, Reducer and Driver classes using MR Testing library.
- Used Maven extensively for building jar files of MapReduce programs and deployed to Cluster.
- Assigned teh tasks of resolving defects found in testing teh new application and existing applications.
- Analyzing teh requirements, designing and developing solutions.
- Managing Project team in achieving teh project goals including resource allocation, resolving technical issues and mentoring teh resources.
- Providing technical solutions/assistance to all development projects
- Advising various teams on teh impact of new enhancements/products
- Bug fixing and 24-7production support running processes.
- Used Linux (Ubuntu) machine for designing, developing and deploying of Java modules
Environment: MapReduce, Pig, Hive, SQOOP, FLUME, HBASE, JDK 1.6, Maven, OS/390, Linux.
Confidential
Python Developer
Responsibilities:
- Utilized standard Python modules such as csv, robot parser, itertools and pickle for development.
- Developed and tested many features for dashboard, created using Bootstrap, CSS, and JavaScript.
- Developed Wrapper in Python for instantiating multi-threaded application.
- Used Python scripts to update content in teh database and manipulate files.
- Generated Python Django Forms to record data of online users.
- Performed troubleshooting, fixed and deployed many Python bug fixes of teh applications and involved in fine tuning of existing processes followed advance patterns and methodologies.
- Skilled in using collections in Python for manipulating and looping through different user defined objects.
- Knowledge on JSON and SimpleJSON based web services.
- Installed numerous python packages using pip and easy install.
- Created independent libraries in Python which can be used by multiple projects which has common functionalities.
- Developed test plan, test scripts and test procedures from teh specification document in Python and automating them to run in teh real time HIL environment.
- Developed and designed automation framework using Python and Shell scripting.
- Developed teh project in Linux environment.
Environment: Python, Perl, MySQL, JavaScript, PHP
