Hadoop Developer/administrator Resume
Berlin, CT
SUMMARY:
- 7+years of versatile experience in analysis, design, development and implementation of software applications and in developing n - tier architecture based solutions with distributed components and internet/ intranet applications.
- Experienced in installing, configuring, and administrating Hadoop cluster of major Hadoop distributions and experience with Hadoop, HDFS, Map Reduce and Hadoop Ecosystem (Pig & Hive), Python
- Expertise in designing and developing applications using J2EE technologies including Servlets, JSP, EJB, JMS, Struts, Hibernate, Web Services, XML, JNDI, JDBC, CVS, Maven, HTML, CSS and JavaScript.
- Experience in Python (libraries used: libraries- Beautiful Soup, numpy, scipy and sublime text, Spyder, pycharm, emacs.
- Hands on experience in writing pig Latin scripts and pig commands.
- Hands on experience in installing, configuring and using ecosystem components like Hadoop Map Reduce, HDFS, Sqoop, Pig, Hive & Python
- Hands-on experience with Hadoop applications (such as administration, configuration management, monitoring, debugging, and performance tuning)
- Configured Zoo Keeper, Cassandra & Flume to the existing Hadoop cluster.
- Extensive experience in all phases of Software Development Life Cycle(SDLC)including identification of business needs and constraints, collection of requirements, detailed design, implementation, testing, deployment and maintenance.
- Expertise in deploying applications using Web and Application Servers like BEA Web Logic 8.1/9.2, IBM Web sphere, Oracle 9iAS, JBoss and Apache Tomcat on Windows and UNIX environments
- Hands on experience with Object Oriented Design (OOD) and developing applications usin g UML. Expertise with Iterative envelopment methodologies, designing Class diagrams, Sequence diagrams and Use case realization.
- Highly proficient in using frame works like Struts, JSF, Hibernate and Design Patterns such as MVC, Sessio n Façade, Front Controller, Data Access Object(DAO), Value Object, Singleton, Service Locator for executing multi-tier, highly scalable, component based, services driven Enterprise Java applications.
- Experience working with Hadoop clusters using Cloudera, HortonWorks, AWS distributions.
- Strong experience with Oracle database and programming languages SQL, PL/SQ Land in developing Packages, Stored Procedures, Functions, Triggers and Cursors.
- Creative problem solver with the ability to rapidly analyze challenges, applying strategic thinking to tactical concerns with strong problem solving skills and a result oriented attitude. Excellent team Player with good Technical, Analytical and interpersonal skills.
PROFESSIONAL EXPERIENCE:
Confidential, BERLIN, CT
HADOOP DEVELOPER/ADMINISTRATOR
RESPONSIBILITIES:
- Involved in gathering the requirements, designing, development and testing
- Involved in designing and implementation of service layer over HBase database.
- Assisted in designing, development and architecture of Hadoop and HBase systems.
- Coordinated with technical teams for installation of Hadoop and third related applications on systems.
- Formulated procedures for planning and execution of system upgrades for all existing Hadoop clusters.
- Supported technical team members for automation, installation and configuration tasks.
- Suggested improvement processes for all process automation scripts and tasks.
- Provided technical assistance for configuration, administration and monitoring of Hadoop clusters.
- Conducted detailed analysis of system and application architecture components as per functional requirements.
- Developed analytical components using Scala, Spark and Spark stream .
- Created scripts to form EC2 clusters for training and for processing.
- Researching and implementing algorithms for large power-law skewed datasets.
- Experience on YARN environment with Storm and Spark
- Developing interactive graph visualization tool based on Prefuse Vis package.
- Developing machine-learning capability via Apache Mahout.
- Hands on experience on Cloud Computing infrastructure like Amazon EMR, S3 and EC2 instance.
- Leading development of terabyte-scale network packet analysis capability.
- Leading research effort to tightly integrate Hadoop and HPC systems.
- Strong Experience in Internet Technologies with experience in J2EE, JSP, Struts, Servlets, JDBC, Apache Tomcat, JBoss, WebLogic, WebSphere, SOAP Protocol, XML, XSL, XSLT, Java Beans, HTML.
- Purchased, deployed, and administered 70 node Hadoop cluster. Administered two smaller clusters.
- Compared Hadoop to commercial big-data appliances from Netezza, XtremeData, and LexisNexis. Published and presented results.
- Wrote the Map Reduce jobs to parse the web logs which are stored in HDFS.
- Developed the services to run the Map-Reduce jobs as per the requirement basis.
- Developing design documents considering all possible approaches and identifying best of them.
- Importing and exporting data into HDFS and HIVE using Sqoop
- Responsible to manage data coming from different sources
- Monitoring the running Map Reduce programs on the cluster.
- Responsible for loading data from UNIX file systems to HDFS. Installed and configured Hive and also written Hive UDFs.
- Involved increasing Hive Tables, loading with data and writing Hive queries which will invoke and run Map Reduce jobs in the backend.
- Implemented the workflows using Apache Oozie framework to automate tasks.
- Developed scripts and automated data management from end to end and sync up b/w all the clusters.
ENVIRONMENT:: Apache Hadoop, Java (jdk1.6), Cassandra, Flatfiles, Oracle11g/10g, MySQL, Toad9.6, Windows NT, UNIX, Sqoop, Hive, Oozie, Python.
Confidential, DURHAM, NCHADOOP DEVELOPER
RESPONSIBILITIES:
- Worked on the proof-of-concept for Apache Hadoop framework initiation.
- Wrote python routines to log into the websites and fetch data for selected options.
- Used python modules such as requests, ur llib, urllib2 for web crawling.
- Used other packages such as Beautiful soup for data parsing.
- Responsible for building scalable distributed data solutions using Hadoop.
- Worked on Installed and configured Hadoop 0.22.0 Map Reduce, HDFS, developed multiple Map Reduce jobs in java for data cleaning and preprocessing.
- Importing and exporting data into HDFS and HIVE using Sqoop
- Experience with spark for streaming data analysis.
- Used Tomcat as application server and Jetty as Servlet containers
- Responsible to manage data coming from different sources
- Monitoring the running Map Reduce programs on the cluster.
- Responsible for loading data from UNIX file systems to HDFS. Installed and configured Hive and also written Hive UDFs.
- Involved increasing Hive Tables, loading with data and writing Hive queries which will invoke and run Map Reduce jobs in the backend.
- Implemented the workflows using Apache Oozie framework to automate tasks.
- Developed scripts and automated data management from end to end and sync up b/w all the clusters.
ENVIRONMENT:: Apache Hadoop, Java (jdk1.6), Cassandra, Jetty, Flatfiles, Oracle11g/10g, MySQL, Toad9.6, Windows NT, UNIX, Sqoop, Hive, Oozie.
Confidential, NEWARK NJHADOOP DEVELOPER
RESPONSIBILITIES:
- Gathered the business requirements from the Business Partners and Subject Matter Experts
- Involved in installing Hadoop Ecosystem components
- Responsible to manage data coming from different sources and Involved in HDFS maintenance and loading of structured and unstructured data
- Worked on analyzing Hadoop cluster and different big data analytic tools including Pig, HBase database and Sqoop.
- Involved in loading data from LINUX file system to HDFS.
- Experience in managing and reviewing Hadoop log files.
- Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
- Importing and exporting data into HDFS and Hive using Sqoop.
- Implemented test scripts to support test driven development and continuous integration.
- Supported in setting up QA environment and updating configurations for implementing scripts with Pig and Sqoop.
- Analyzed large data sets by running Hive queries and Pig scripts.
- Worked on tuning the performance Pig queries.
- Mentored analyst and test team for writing Hive Queries.
- Installed Oozie workflow engine to run multiple Map Reduce jobs.
- Worked with application teams to install operating system, Hadoop updates, patches, version upgrades as required.
ENVIRONMENT:: Hadoop, HDFS, Map Reduce, Hive, Pig, Sqoop, Linux, Java, Oozie, HBase
Confidential, CHERRY HILL, NJJAVA/J2EEDEVELOPER
RESPONSIBILITIES:
- Responsible for design and development using J2EE architecture.
- Developed the presentation layer for the credit enhancement module in JSP Struts1.2 was used to implement the Model View Layer (MVC) architecture. Client Side validations were done using JavaScript.
- Used AJAX for performing validations in the credit enhancement module.
- Developed the global logging module which was used across all the modules usi ng Log4J component
- Development of persistent components using Hibernate 3.0.
- Used JSF for UI component representation and basic server side validations.
- Used Web services to expose the business logic
- Worked with Web sphere ESB mediation modules flow components Routing, Database lookup, Database logging and Structure transformation
- Written SQL queries, stored procedures and modifications to existing database structure.
- Developed complex PL/SQL stored procedures, functions and triggers
- Responsible for delivering fail-safe, scalable, clustered application system with load-bala ncer.
- Used replicated database to provide safety of data and scalability of the system.
- Applied Agile methodologies for software development
- Involved in production support and maintaining the application after production.
ENVIRONMENT:: Java1.4, J2EE1.4, JSP1.2, Servlets, Struts1.2, Hibernate3.0, HTML, JSF, JavaScript, SQL, PL/SQL, Oracle 9i,Python, Web logic 8.1 SP6, Windows NT, MVS, Eclipse EE, XML, CVS, Rational Clear Quest.
Confidential, MANHATTAN, NYCJAVADEVELOPER
RESPONSIBILITIES:
- Developed JSP pages, Servlets and HTML pages as per requirement.
- Coding using Java, Java Script and HTML .
- Developed the necessary Java Beans, PL/SQL procedures for the implementation of business rules.
- Designed, developed and executed Data Migration from Db2 Database to Oracle Database using Linux scripts, Java and SQL loader concepts.
- Developed UNIX and java utilities for Data migration from Db2to Oracle . Sole developer and SPOC for the migration Activity.
- Used JDBC to provide database connectivity to database tables in Oracle.
- Developed the web interface using JSP and developed struts action classes.
- Used Web Sphere Application Server for application deployment.
ENVIRONMENT:: JDK5.0, J2EE, IBMDB2, IBM Web Sphere Application Server, IBMWSAD, Struts1.x, EJB, JSP, Servlets, HTML, CVS, JavaScript, Oracle 9i, UNIX Scripting and Windows 2000.
