Python Developer Resume
Chicago, IL
SUMMARY
- Over 6 years IT experience with strong emphasis on requirement gathering, analysis, design, development, implementation, testing and development of software applications in Hadoop, HDFS, MapReduce, Hadoop Ecosystem and RDBMS.
- Experience in Big data Hadoop,Hadoop Ecosystem components like MapReduce, Sqoop, Flume, Kafka, Pig,Hive, Spark, Storm, HBase, Airflow, Oozie, and Zookeeper.
- Worked extensively on installing and configuringHadoopecosystem components Hive, SQOOP, HBase, Flumeand Zookeeper.
- Good experience of software development in Python (libraries used: Beautiful Soup, NumPy, SciPy, Pandas dataframe, Matplotlib, network, urllib2, MySQL dB for database connectivity) and IDEs - sublime text, Spyder, PyCharm, Visual Studio Code.
- Expert in using Django Authentication system, Django templating system, creating models and forms.
- Hand full experience on AWS services like S3 bucket, lambda, API Gateway, Cognito, Step Function, RDS and Cloud Watch using CDK.
- Created network architecture on AWS VPC, subnets, Internet Gateway, Route. Perform S3 buckets creation, configured the storage on S3 buckets, policies and the IAM role-based policies.
- Experience implementing Cloud based Linux OS in GCP to Develop Scalable Applications with Python.
- Hands of experience in GCP, Big Query, GCS bucket, G - cloud function, cloud dataflow, Pub/suB cloud shell, GSUTIL, BQ command line utilities, Data Proc, Stack driver.
- Good experience in developing web applications implementing MVT/MVC architecture using Django, Flask and spring web application frameworks.
- Experience with Oozie Workflow Engine in running workflow jobs with actions dat run Hadoop Map/Reduce and Pig jobs.
- Experience in understanding the security requirements for Hadoop and integrate with Kerberos authentication and authorization infrastructure.
- Experience in configuring the Zookeeper to coordinate the servers in clusters and to maintain the data consistency.
- Acumen on Data Migration from Relational Database to Hadoop Platform using Sqoop.
- Experienced in Data Analysis, Data mining, Acquisition, Validation, Visualization and discovering meaningful business insights on large data sets of Structured and Unstructured data.
- Hands on experience in developing Spark applications using Spark tools like RDD transformations, Spark core, Spark Streaming and Spark SQL.
- Implemented Sqoop for large dataset transfer between Hadoop and RDBMs.
- Experience setting up instances behind Elastic Load Balancer in AWS for high availability.
TECHNICAL SKILLS
Primary Languages: Python 3.x,2.7/2.4, SQL
Python Framework: Django, Flask
Python Libraries: Pandas, NumPy, SciPy, MatPlotlib, urllib2, Networkx, MySQLdb
IDE Tool: Pycharm, Eclipse, PyDev
Web Technology: JavaScript, JQuery, HTML, CSS, Bootstrap, Angular.js
RDBMS: MS-SQL, MySQL, Oracle
Operating Systems: Windows, LINUX
PROFESSIONAL EXPERIENCE:
Confidential, Chicago, IL
Python Developer
Responsibilities:
- Developed Architecture for Parsing applications to fetch the data from different services and transforming to store in different formats.
- Developed parsers for Extracting data from different sources of web services and transforming to store in various formats such as CSV, Database files, HDFS storage, etc. then to perform analysis.
- Parsers written in Python for extracting useful data from the design data base. Used Parsekit (Enigma.io) framework for writing Parsers for ETL extraction.
- Implemented Algorithms for Data Analysis from Cluster of Web services.
- Worked with lxml to dynamically generate SOAP requests based on the services. Developed custom Hash-Key (HMAC) based algorithm in Python for Web Service authentication.
- Worked with Report Lab PDF library to dynamically generate the PDF documents with Images and data retrieved from various sources of Web services.
- Built the Web API on the top of Django framework to perform REST methods. Used MongoDB and MySQL databases in Web API development. Developed database migrations using SQLAlchemy Migration.
- Generated graphical reports using python package NumPy and matplotlib.
- Usage of advance features like pickle/unpickle in python for sharing the information across the applications.
- Managed datasets using Panda data frames and MySQL, queried MYSQL database queries from python using Python-MySQL connector and MySQL db package to retrieve information.
- Utilized Python libraries wxPython, NumPy, Twisted and matplotlib.
- Wrote Python scripts to parse XML documents and load the data in database.
- Used Wireshark, live http headers, and Fiddler2 debugging proxy to debug the Flash object and help the developer create a functional component. The PHP page for displaying the data uses AJAX to sort and display the data. The page also outputs data to .csv for viewing in Microsoft Excel.
- Added support for Amazon AWSS3 and RDS to host static/media files and the database into Amazon Cloud.
- Writing Python scripts with Cloud Formation templates to automate installation of Auto scaling, EC2, VPC, and other services.
- Used Docker containers for development and deployment.
- Familiar with UNIX / Linux internals, basic cryptography & security.
- Developed multiple spark batch jobs in Scala using Spark SQL and performed transformations using many APIs and update master data in Cassandra database as per the business requirement.
- Written Spark-Scala scripts, by creating multiple UDF's, spark context, Cassandra sql context, multiple API's, methods which support data frames, RDD's, data frame Joins, Cassandra table joins and finally write/save the data frames/RDD's to Cassandra database.
- As part of the POC migrated the data from source systems to another environment using Spark, SparkSQL.
- Developed and implemented core API services using Python with spark.
- Created data frames schema from raw data stored at Amazon S3 using PySpark.
- Used PySpark Data frame for creation of table and performing analytics over it.
- Using Jenkins AWS Code Deploy plugin to deploy to AWS.
- Developed tools using Python, XML to automate some of the menial tasks. Interfacing with supervisors, artists, systems administrators, and production to ensure production deadlines are met.
Environment: Python 3.x, Parse kit (Enigma.io), Django, Flask, lxml, SUDS, HMAC, pandas, Numpy, matplotlib, MongoDB, MySQL, SOAP, REST, PyCharm, Docker, AWS (EC2, S3).
Confidential, Boston, MA
Hadoop Developer
Responsibilities:
- Worked closely with the business analysts to convert the Business Requirements into Technical Requirements and preparing low and high-level documentation.
- Performing transformations using Hive, MapReduce, hands on experience in copying .log, snappy files into HDFS from Greenplum using Flume & Kafka, loaded data into HDFS and extracted the data into HDFS from MYSQL using Sqoop.
- Imported required tables from RDBMS to HDFS using Sqoop and used Storm/ Spark streaming and Kafka to get real time streaming of data into HBase.
- Experience in building multiple Data pipelines, end to end ETL and ELT process for Data ingestion and transformation in GCP.
- Developed views and templates with Python and Django's view controller and templating language to create a user-friendly website interface.
- Experience in Writing Map Reduce jobs for text mining and worked with predictive analysis team and Experience in working with Hadoop components such as HBase, Spark, Yarn, Kafka, Zookeeper, PIG, HIVE, Sqoop, Oozie, Impala and Flume.
- Wrote HIVE UDF's as per requirements and to handle different schema’s and xml data.
- Implemented ETL code to load data from multiple sources into HDFS using Pig Scripts.
- Developed data pipeline using Python, hive to load data into data link. Perform data analysis data mapping for several data sources.
- Designed new Member and Provider booking system which allows providers to book new slots, with sending out the member leg and provider Leg directly to TP through Datalink.
- Write a Python program to maintain raw file archival in GCS bucket.
- Open SSH tunnel to Google DataProc to access to yarn manager to monitor spark jobs.
- Analyze various type of raw file like Json, Csv, Xml with Python using Pandas, NumPy etc.
- DevelopedSparkapplications using Scalafor easy Hadoop transitions.And Hands on experienced in writingSparkjobs and Sparkstreaming API using Scalaand Python.
- Used SparkAPI over Cloudera Hadoop YARN to perform analytics on data in Hive developedSpark code andSpark-SQL/Streaming for faster testing and processing of data.
- Installed Oozie workflow engine to run multiple Hive and Pig jobs.
- Designed and developed User Defined Function (UDF) for Hive and Developed the Pig UDF’S to pre-process the data for analysis as well as experience in (UDAFs) for custom data specific processing.
- Created Airflow Scheduling scripts in Python.
- Automated the existing scripts for performance calculations using scheduling tools like Airflow.
- Designed and developed the core data pipeline code, involving work in Java and Python and built onKafkaand Storm.
- Good noledge on Partitions, bucketing concepts in Hive and designed both Managed and External tables in Hive for optimized performance.
- Performance tuning using Partitioning, bucketing of IMPALA tables.
- Created cloud-based software solutions written in Scala Spray IO, Akka, and Slick.
- Hands on experience on fetching the live stream data from DB2 to HBase table using Spark Streaming and Apache Kafka.
- Experience in job workflow scheduling and monitoring tools like Oozie and Zookeeper.
- Worked on NoSQL databases including HBase and Cassandra.
- Populated HDFS and Cassandra with huge amounts of data using Apache Kafka.
Environment: Map Reduce, HDFS, Hive, Pig, HBase, Python, SQL, Sqoop, Flume, Oozie, Impala, Scala, Spark, Apache Kafka, Play, GCP, AKKA, Zookeeper, J2EE, Linux Red Hat, HP-ALM, Eclipse, Cassandra, SSIS.
Confidential, San Diego, CA
Hadoop/Python Developer
Responsibilities:
- Worked extensively on Hadoop Components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, YARN, Spark and Map Reduce programming
- Converting the existing relational database model toHadoopecosystem.
- Worked with Linux systems and RDBMS database on a regular basis in order to ingest data using Sqoop.
- Developed Schedulers dat communicated with the Cloud based services (AWS) to retrieve the data.
- Strong experience in working with Elastic MapReduce and setting up environments on Amazon AWS EC2 instances.
- Ability to spin up different AWS instances including EC2-classic and EC2-VPC using cloud formation templates.
- Collected data using Spark Streaming from AWS S3 bucket in near-real-time and performs necessary Transformations and Aggregations to build the data model and persists the data in HDFS.
- Imported the data from different sources like AWS S3, LFS into Spark RDD.
- Experienced in working with Amazon Web Services (AWS) EC2 and S3 in Spark RDD.
- Managed and reviewed Hadoop and HBase log files.
- Worked extensively with importing metadata into Hive and migrated existing tables and applications to work on Hive and AWS cloud.
- Designed and implementedHIVE queries and functions for evaluation, filtering, loading and storing of data.
- Analyze table data and implement compression techniques like Teradata Multivalued compression.
- Involved in ETL process from design, development, testing and migration to production environments.
- Involved in writing the ETL test scripts and guided the testing team in executing the test scripts.
- Involved in performance tuning of the ETL process by addressing various performance issues at the extraction and transformation stages.
- Writing Hadoop MapReduce jobs to run on Amazon EMR clusters and creating workflows for running jobs.
- Generating analytics reporting on probe data by writing EMR (elastic map reduce) jobs to run on Amazon VPC cluster and using Amazon data pipelines for automation.
- Worked with Elastic MapReduce (EMR) on Amazon Web Services (AWS).
- Have good understanding of Teradata MPP architecture such as Partitioning, Primary Indexes.
- Good noledge in Teradata Unity, Teradata Data Mover, OS PDE Kernel internals, Backup and Recovery.
- Created HBase tables to store variable data formats of data coming from different portfolios.
- Created Partitions, Buckets based on State to further process using Bucket based Hive joins.
- Involved in transforming data from Mainframe tables to HDFS, and HBase tables using Sqoop.
- Creating Hive tables and working on them using HiveQL.
- Creating and truncating HBase tables in hue and taking backup of submitter ID.
- Developed data pipeline using Kafka to store data into HDFS.
- Used Spark API overHadoopYARN as execution engine for data analytics using Hive.
- Continuous monitoring and managing theHadoop cluster through Cloudera Manager.
- DevelopedETLProcess usingHIVE and HBASE.
- Prepared the Technical Specification document for the ETL job development.
- Responsible to manage data coming from different sources.
- Loaded the CDRs from relational DB using Sqoop and other sources toHadoop cluster by using Flume.
- Experience in processing large volume of data and skills in parallel execution of process using Talend functionality.
- Installed and configured Apache Hadoop, Hive and Pig environment.
Environment: Hadoop, HDFS, pig, Hive, Flume, Sqoop, Oozie, Python, Shell Scripting, SQL Talend, Spark, HBase, Elastic search, Linux- Ubuntu, Kafka.
Confidential
Software Engineer
Responsibilities:
- Involved in the analysis, design, implementation, and testing of the project.
- Strong understanding and practical experience in developing Spark applications with Python.
- Developed Scala scripts, UDFs using both Data frames/SQL in Spark for Data Aggregations.
- Designed, develop, test, deploy and maintain the website.
- Developed entire frontend and backend modules using Python on Django Web Framework.
- Developed Python scripts to update content in the database and manipulate files.
- Rewrite existing Java application in Python module to deliver certain format of data
- Developed entire frontend and backend modules using Python on Django Web Framework.
- Generated property list for every application dynamically usingPython.
- Designed and developed the UI of the website using HTML, XHTML, AJAX, CSS and JavaScript.
- Wrote Python scripts to parse XML documents and load the data in database.
- Generated property list for every application dynamically using Python.
- Handled all the client-side validation using JavaScript.
- Performed testing using Django’s Test Module.
- Designed and developed data management system using MySQL.
- Creating unit test/regression test framework for working/new code.
- Responsible for search engine optimization to improve the visibility of the website.
- Responsible for debugging and troubleshooting the web application.
Environment: Python, Django, Java, MySQL, Linux, HTML, XHTML, CSS, AJAX, JavaScript, Apache Web Server.
